Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Where a protein goes

A ribosome makes every protein in the same cytosol, and a digestive enzyme, a histone and a respiratory complex all have to end up somewhere else.

The previous lesson decided which proteins get made. This one follows a finished chain to its destination, which turns out to be a problem the cell solves with address labels written into the protein itself. The lesson ends with the two compartments whose addressing systems look nothing like the others, for a reason that has to do with where they came from.

Watching the route

George Palade, at the Rockefeller Institute through the 1960s, established the secretory pathway with a technique that made a chemical process visible in space and time. The tissue was guinea pig pancreas, which secretes digestive enzymes and is therefore almost entirely devoted to the process being studied.

The method is pulse-chase autoradiography. Give the tissue radioactive leucine for three minutes, then flood it with unlabelled leucine so no further label is incorporated, and fix samples at intervals. Cut thin sections, coat them with photographic emulsion, and the silver grains that develop mark where the labelled protein was at that moment.

The sequence is unambiguous. At three minutes the label is over the rough endoplasmic reticulum. By about seven minutes it has moved to the transitional region and the Golgi. By thirty to forty minutes it is in condensing vacuoles on the far side of the Golgi, and after an hour or two it is in mature zymogen granules waiting at the cell surface for the signal to release. Palade shared a Nobel Prize in 1974 with Albert Claude and Christian de Duve.

Two conclusions follow immediately. Secretion is a directional route through a series of distinct compartments rather than diffusion to the surface. And a protein destined for secretion enters the endoplasmic reticulum at the very beginning, while it is still being made, which raises the question the next section answers: how does a ribosome know?

The signal hypothesis

Günter Blobel and David Sabatini proposed in 1971 that a secreted protein carries a short sequence at its own front end that directs the ribosome making it to the endoplasmic reticulum membrane. The proposal was speculative and it was tested in 1975 by Blobel and Bernhard Dobberstein in a cell-free system, which is the cleanest kind of demonstration available.

Translate the messenger RNA for a secretory protein in vitro with no membranes present, and the product is slightly longer than the mature protein isolated from tissue, carrying an extra stretch at its amino end. Translate the same message with microsomal membranes, small vesicles derived from the endoplasmic reticulum, and the product is the correct length, is inside the vesicles, and is protected from added protease. Add the membranes after translation is finished and nothing happens: the protein stays outside and stays long.

Every element of the hypothesis is confirmed by that pattern. There is an extra sequence, it is removed during transport, transport happens across the membrane rather than around it, and it must occur while the chain is still being made. Blobel received a Nobel Prize in 1999.

The mechanism as now understood: the signal sequence, typically 15 to 30 residues with a hydrophobic core, emerges from the ribosome and is bound by the signal recognition particle, which is itself a complex of RNA and protein. Binding pauses translation, which prevents the chain from being made in the wrong place. The particle docks with its receptor on the endoplasmic reticulum, hands the ribosome to a protein-conducting channel called the translocon, and translation resumes with the growing chain threading directly into the channel. Signal peptidase clips the signal off on the far side.

Note how the pause solves a real problem. Without it, a chain would keep growing and might fold in the cytosol before reaching the membrane, and a folded protein cannot be threaded through a narrow channel.

Example. Some proteins reach the endoplasmic reticulum after translation is complete rather than during it, and yeast uses this route extensively. What extra machinery must such a route require, and what does its existence say about the signal hypothesis?

It must require something to keep the finished chain unfolded, or to unfold it again, because the translocon passes an extended chain and not a folded protein. Post-translational translocation accordingly depends on cytosolic chaperones that hold the protein in a loosely folded state, and on a motor on the far side, a chaperone of the Hsp70 family in the endoplasmic reticulum lumen, that binds the emerging chain and ratchets it through, since without the ribosome pushing there is nothing to drive the direction. What its existence says about the signal hypothesis is that the address and the timing are separable. The signal sequence specifies the destination; whether the chain travels during or after synthesis is a separate question answered by the hydrophobicity of the signal and by which machinery binds it first. That separation is worth having in mind generally, because the same logic applies to mitochondrial and chloroplast import, which are entirely post-translational and use unfolding chaperones for the same reason.

Now you. A protein made with no recognisable targeting sequence at all ends up in the cytosol. Is that a mechanism, and how would you test your answer?

Answer

It is the absence of a mechanism, and that is the point: the cytosol is the default destination, because it is where synthesis happens and nothing has to move for a protein to stay there. Every other compartment requires a positive signal. The test is a transplantation experiment and it has been done many times: take a protein that is normally cytosolic, attach a signal sequence from a secretory protein to its front end, and see whether it is now secreted. It is. Do the converse, delete the signal from a secretory protein, and it accumulates in the cytosol. Fusing a targeting sequence to a reporter such as green fluorescent protein and watching where the fluorescence goes is the routine modern version, and it establishes both that the sequence is sufficient and, with the deletion, that it is necessary. That pair of experiments, sufficiency and necessity, is the standard form of an argument that a sequence is an address rather than a correlate, and it is worth asking of any claim that some sequence "targets" something.

The other addresses

Each compartment has its own signal and its own machinery, and the pattern of differences is not arbitrary.

A nuclear localisation signal is a short run of basic residues, the classic example being the five lysines and arginines of the SV40 large T antigen. It is not cleaved after use, because a protein in the nucleus must be reimported after every mitosis, when the nuclear envelope breaks down and reforms. Importins carry the cargo through the nuclear pore complex, an enormous assembly of around 120 megadaltons. Molecules under roughly 40 kilodaltons diffuse through the pore passively; larger ones need a signal and a carrier, and the directionality is set by a gradient of a small GTPase, Ran, maintained across the envelope.

A mitochondrial matrix targeting sequence is an amphipathic helix, positively charged along one face, at the amino end, and it is cleaved on arrival. Import is post-translational, requires the protein to be unfolded, and, crucially, requires the membrane potential across the inner membrane, since the positively charged presequence is drawn electrophoretically into a negative matrix. That makes protein import one of the many things a mitochondrion cannot do once its gradient collapses.

A peroxisomal signal is usually just three residues at the far, carboxy end, most often serine-lysine-leucine. Peroxisomes are unusual in importing folded proteins, even oligomers, which requires a transport mechanism quite unlike a narrow channel and is still not fully explained.

The pattern is that where a protein must remain identifiable for reimport the signal is kept, and where it is used once it is cut off; where the chain must be threaded the signal is at the front, and where it is recognised on a finished protein it may be anywhere.

Folding, and why sequence is not quite enough

Christian Anfinsen showed in the early 1960s that ribonuclease, denatured completely and then returned to normal conditions, refolds spontaneously to full activity. The information for the three-dimensional structure is in the sequence, and no external instruction is needed. He received a Nobel Prize in 1972.

That result is true and incomplete, for two reasons.

The first is Cyrus Levinthal's observation of 1969. A chain of 100 residues with only three possible conformations per residue has 3100=5×1047 possible structures, and sampling them at 10-13 s each would take 1.6×1027 years, seventeen orders of magnitude longer than the age of the universe. Proteins fold in milliseconds to seconds. So folding is not a search: it is a directed process down an energy landscape shaped like a funnel, in which local structure forms first and constrains what follows.

The second is concentration. Anfinsen's experiment used dilute pure protein. A cell's cytosol carries protein at 200 to 300 grams per litre, and a partly folded chain has hydrophobic surfaces exposed that would rather stick to a neighbour than to itself. Aggregation is a competing reaction and at cellular concentrations it is fast.

Chaperones exist to lose that race. Hsp70 proteins bind short exposed hydrophobic stretches and release them in an ATP-driven cycle, giving the chain repeated chances to fold without staying exposed long enough to aggregate. The chaperonins, GroEL with its cap GroES in bacteria, go further and provide a barrel-shaped cavity in which a single protein of up to about 60 kilodaltons is enclosed and allowed to fold in isolation, physically prevented from meeting a partner.

The important honest point is that chaperones do not tell a protein what shape to adopt. They raise the yield of a reaction whose endpoint the sequence already determines, and Anfinsen's conclusion survives. Every one of them is called a heat shock protein because they were found as the proteins induced by high temperature, which is exactly the condition that unfolds things.

Destruction on schedule

A cell controls protein levels at both ends, and degradation is as regulated as synthesis. Protein half-lives in a cell range from a couple of minutes to several days.

The major route is the ubiquitin-proteasome system, worked out by Aaron Ciechanover, Avram Hershko and Irwin Rose, who shared a Nobel Prize in 2004. A small protein, ubiquitin, is attached to a lysine of the target through a cascade of three enzyme activities, and further ubiquitins are added to make a chain. The chain is recognised by the proteasome, a barrel-shaped complex that unfolds the substrate, feeds it into an interior chamber, and cuts it into short peptides, spending ATP throughout.

The specificity lies in the third enzyme of the cascade, the ubiquitin ligase, and humans have more than six hundred of them. That is how a general destruction machine acquires selective targets: the machine is common and the labellers are many. The signals recognised range from a specific sequence exposed only when a protein is damaged, to a phosphate added by a kinase in response to a hormone, to the identity of the amino-terminal residue itself.

Quality control uses the same machinery. A protein that fails to fold in the endoplasmic reticulum is retained, retro-translocated back into the cytosol and degraded, a process called ERAD, and a backlog of unfolded protein triggers a signalling programme, the unfolded protein response, that slows translation and increases chaperone production.

This is not an abstract cleanup service. The commonest mutation causing cystic fibrosis, deletion of a single phenylalanine at position 508 of the CFTR chloride channel, does not destroy the channel's function. The mutant protein, if it reaches the cell surface, works substantially. What it fails to do is fold quickly enough to pass quality control, so it is caught in the endoplasmic reticulum and destroyed, and the cell surface has no channel at all. The disease is a trafficking failure, which is why one class of drug for it, the correctors, works by helping the protein fold and escape rather than by fixing the channel.

Example. In I-cell disease the enzymes that belong in lysosomes are found instead at high concentration in the patient's blood, and their cells accumulate undigested material. One enzyme is defective, and it is not any of the lysosomal enzymes. What kind of enzyme must it be?

It must be the one that writes the address. Lysosomal enzymes are tagged in the Golgi with mannose 6-phosphate, added in two steps of which the first is carried out by a phosphotransferase, and a receptor then recognises that tag and diverts the tagged proteins into vesicles bound for the lysosome. Lose the phosphotransferase and every lysosomal enzyme is made normally, folds normally and is fully active, but carries no tag, so it is not diverted and follows the default secretory route to the outside of the cell. Hence enzymes in the blood at up to twenty times normal levels and none where they are needed. The reasoning generalises to a useful diagnostic principle: when many unrelated proteins are simultaneously in the wrong place, suspect the addressing machinery rather than the proteins, and when one is, suspect that protein's own signal. It is also a clean demonstration that a sorting signal can be a sugar rather than a peptide sequence, which nothing in the signal hypothesis required.

Now you. Lysosomal enzymes work at pH 5, maintained by a proton pump in the lysosomal membrane, while the cytosol sits near pH 7.2. What safety property does that give a cell, and what does it imply about the pump?

Answer

A lysosome is full of enzymes that would digest the cell, so the pH difference is a containment mechanism as well as a working condition. An enzyme with a sharp optimum at pH 5 is largely inactive two pH units higher, so a lysosome that leaks a little releases proteases and lipases into a compartment where they do comparatively little damage. Containment therefore does not rely on the membrane being perfect, which no membrane is. It implies that the pump is essential and continuously active, since protons leak back and the gradient would otherwise dissipate, and it means an inhibitor of that pump should stop lysosomal digestion without touching any lysosomal enzyme. That is exactly what bafilomycin does in the laboratory, and it is why chloroquine, a weak base that accumulates in acidic compartments and raises their pH, interferes with lysosomal function. The pattern, a compartment defined by a gradient that a pump maintains, is the same one this course established for the plasma membrane and for the mitochondrion, used here for a third purpose.

Two compartments that came from somewhere else

Mitochondria import over a thousand nuclear-encoded proteins by the route described above. They also make thirteen of their own, from their own genome, on their own ribosomes.

The list of oddities is long and it is coherent. A human mitochondrion carries a circular DNA of 16,569 base pairs with 37 genes: 13 proteins, 22 transfer RNAs and 2 ribosomal RNAs. Its ribosomes are bacterial in type rather than eukaryotic, and are inhibited by antibiotics such as chloramphenicol that do not touch the cell's own. Its protein synthesis starts with formylmethionine, as bacteria do and as the cytosol does not. It has two membranes, and the inner one contains cardiolipin, a lipid otherwise characteristic of bacteria. It reads a slightly different genetic code. And no cell makes a mitochondrion from scratch: they arise only by the division of existing mitochondria, which is Virchow's principle applied one level down.

Lynn Margulis argued in 1967, against considerable resistance, that the explanation is endosymbiosis: a bacterium taken up by an ancestral cell and retained. Molecular phylogenetics has since placed the mitochondrial genes firmly within the alphaproteobacteria and the chloroplast genes within the cyanobacteria, and the case is now as settled as anything in cell biology.

Two questions remain live and are worth stating honestly. Why has almost the entire genome moved to the nucleus, and why has any of it stayed? The transfer is easy to motivate: a gene in the nucleus is protected by a nuclear envelope, repaired by the full nuclear repair machinery, and inherited by orderly meiosis, while a gene in a mitochondrion sits next to the most oxidising chemistry in the cell. Human mitochondrial DNA mutates roughly ten to twenty times faster than nuclear DNA.

Why thirteen genes remain is less settled. The retained proteins are among the most hydrophobic in the cell, core subunits of the respiratory complexes, and it may simply be that importing them across two membranes is not feasible. A second proposal, John Allen's co-location for redox regulation hypothesis, is that these particular subunits must be made under the direct local control of the redox state of the chain they belong to, which requires their genes to be present in the same compartment. The two are not exclusive and the question is open.

The consequences reach into medicine. Mitochondria are inherited maternally, since the sperm contributes essentially none, so mitochondrial diseases show a distinctive pedigree: an affected mother passes it to all her children, an affected father to none. A cell carries hundreds or thousands of mitochondrial genomes, so a mutation is usually present in some fraction of them, called heteroplasmy, and symptoms appear only above a threshold fraction that differs by tissue. That is why mitochondrial disorders such as MELAS and Leber's optic neuropathy vary so widely in severity between relatives carrying the same mutation, and why tissues with the highest energy demand, nerve and muscle, are affected first.

Example. Chloramphenicol inhibits bacterial ribosomes and was widely used until its side effects limited it. One of those side effects is suppression of the bone marrow. Explain it, and predict which other tissues should be vulnerable.

Mitochondrial ribosomes are bacterial in type, so a drug selected for inhibiting bacterial translation will also inhibit the synthesis of the thirteen mitochondrially encoded proteins, all of which are core subunits of the respiratory chain and ATP synthase. Cells cannot make good the loss from the nucleus, so respiratory capacity falls. The tissues that should suffer first are those that divide fastest or demand the most ATP, which is precisely the bone marrow, where blood cells are produced continuously and in enormous numbers. The prediction extends correctly to other drugs: linezolid, a modern antibiotic that also targets a bacterial ribosome site, causes marrow suppression, peripheral and optic neuropathy, and lactic acidosis on prolonged use, and lactic acidosis is exactly what a partial block of oxidative phosphorylation should cause, since the cell falls back on the glycolysis and fermentation route described earlier in this course. This is a case where an evolutionary fact, the ancestry of an organelle, predicts a drug's toxicity profile, and it is a good answer to anyone who asks what endosymbiosis is good for knowing.

Now you. Some parasitic protists, including Giardia and the microsporidia, were once thought to have no mitochondria and to represent lineages that split before the endosymbiosis. That interpretation has collapsed. What would you look for to test it, and what do you think was found?

Answer

Look for two things: nuclear genes of clear alphaproteobacterial ancestry that encode mitochondrial proteins, and a residual double-membrane organelle that no longer respires. Both were found. These organisms carry nuclear genes for mitochondrial chaperones and iron-sulfur cluster assembly proteins whose phylogeny places them with the mitochondrial lineage, and they possess reduced organelles, mitosomes in Giardia and hydrogenosomes in some other anaerobes, which are double-membraned, are derived from mitochondria, and have lost the respiratory chain while retaining iron-sulfur cluster assembly. So these are not early-branching lineages that missed the endosymbiosis; they are lineages that acquired mitochondria and then reduced them almost to nothing while living anaerobically. Two lessons follow. Absence of a structure is weak evidence for absence of an ancestor, since loss is common and easy, and any argument of the form "this organism lacks X so it diverged before X" needs a genome to support it. And iron-sulfur cluster assembly appears to be the one mitochondrial function that has never been lost in any eukaryote examined, which suggests it, rather than respiration, may be the function that made the organelle indispensable in the first place.

A protein now has an address, a folded structure, a compartment and a scheduled end. What it does not have is a way of getting anywhere quickly. The first lesson of this course computed that diffusion covers a micrometre in a millisecond and a metre in thirty years, and a eukaryotic cell is large enough for that to matter: a secretory vesicle must reach the surface, a mitochondrion must be positioned where ATP is needed, and a chromosome must be dragged to one end of a dividing cell. The next lesson is the internal scaffolding that makes position possible and the motors that walk along it.