Directed evolution of enzymes
A Socratic walk-through of directed evolution of enzymes — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #How can a chemist obtain an enzyme for a reaction no organism ever needed?
Enzymes are the best catalysts we know of, and every one of them was produced by evolution for a purpose some organism had. So when a chemist wants an enzyme that tolerates an industrial solvent, or that forms a bond no living thing has ever needed to form, the obvious route seems blocked. Nature did not make one, and designing a protein from first principles means predicting how a chain of hundreds of amino acids folds and what its folded shape does to a transition state — a calculation that remains hard.
Yet such enzymes exist and are sold. Frances Arnold shared the 2018 Nobel Prize in Chemistry for the method that produces them. What can the method be doing, if it is neither borrowing from nature nor calculating the answer?
Reasoning it through
REASONING #Begin with why the design route is so hard. A modest protein of a hundred residues has twenty choices at each position, giving twenty to the hundredth possible sequences — a number with well over a hundred digits. No search can enumerate that. Any method that works must avoid searching the space at all.
So ask what evolution itself does, since it faces the same arithmetic. It does not search globally. It starts from something that already works, changes it slightly, keeps whatever does better, and repeats. Each round moves a short distance from a functioning protein, so the sequences examined are overwhelmingly ones that still fold. The space is astronomical but the walk is local, and locality is what makes it tractable.
Can that be run deliberately? Three things are needed, and each has a laboratory equivalent. Variation: introduce random mutations with an error-prone copying step, or recombine fragments of related genes. Heredity: the variants are genes, so a winner can be copied and mutated again. And the third — the one the whole method turns on — a rule that decides which variants continue.
Look carefully at that third element, because it is where the answer to the original question lives. The chemist does not know what the improved enzyme should look like; that is the entire difficulty. What the chemist can state is what it should do. So the laboratory supplies not a design but a filter: an assay run on every variant, or better, an arrangement in which cells only survive or grow when their protein performs the desired chemistry. All the information about the target reaction enters through that filter, and none of it through the sequences, which are generated blindly.
This inverts where the intelligence sits. In conventional design the hard thinking goes into the molecule. Here it goes into the test. That is why an enzyme for an unprecedented reaction is possible at all: nature never needed the reaction, but the experimenter can still recognise it when it happens, and recognising it is sufficient.
One structural requirement is easy to overlook. Each protein's performance must stay attached to the gene that encoded it, or the winners cannot be bred. Keeping every variant inside its own cell or droplet does this, which is why libraries are grown in bacteria or compartmentalised rather than mixed in one flask.
And a warning follows directly from the same logic. If the filter is the only channel through which the goal is expressed, then the filter is what you get — Arnold's own summary is that you get what you screen for. Screen at thirty degrees and you may get an enzyme that fails at sixty. Screen on a convenient surrogate substrate and you may get an enzyme excellent on the surrogate and indifferent to the real one. Every mismatch between the assay and the intended use is faithfully converted into a defect in the product.
Does the local walk ever get stuck? Yes. Interactions between residues mean an improvement can require two changes at once, each neutral or harmful alone, so the fitness landscape has ridges a single-step walk cannot cross. The practical answers are evolution's own: let neutral variation accumulate without selecting hard, and recombine successful variants so changes found separately can combine.
The analogy
THE ANALOGY #Imagine breeding a dog for a job no dog has done. You cannot specify the animal — you have no theory of what its bones and temperament should be. What you can do is set the trial, run every pup through it, and breed from whoever comes closest. Ten generations later there is an animal well suited to the job, and nobody ever knew in advance what it should look like. The trial did the designing.
Dog breeding starts from a population that varies naturally and moves slowly, whereas a laboratory manufactures its own variation at a chosen rate and can screen far more offspring per round than a breeder ever sees — so it compresses into weeks what would take a breeder a lifetime.
Clarifying the model
THE MODEL #Two clarifications are worth making.
The first is that "random" applies to the mutations, not to the method. Modern practice is heavily targeted: structural knowledge and sequence comparisons are used to decide which residues to randomise, so the library concentrates variation where it plausibly matters. And design and evolution are increasingly used together — a computationally designed starting protein with weak activity is often a better parent than any natural enzyme, with directed evolution then supplying the many-fold improvement that design alone does not reach.
The second is what "no organism ever needed" really means. The evolved catalysts that form carbon-silicon bonds, or transfer carbenes, are not built from new chemistry invented in the lab; they are natural proteins, often heme-containing ones such as cytochromes, whose existing reactive centre has been coaxed into a reaction it was capable of only weakly and incidentally. Directed evolution amplifies a latent, promiscuous side activity into a main one. The method reliably improves something already present at a low level, and is far less able to conjure activity where there is none to select on. Getting a measurable starting signal is usually the hardest part of a campaign.
A picture of it
THE PICTURE #How to readOne round of a campaign, in illustrative proportions. The band on the left is the library of gene variants made from one parent; the three on the right are what the assay sorts them into. The width of the top two is the point: nearly everything made is wasted, and the method tolerates that because variants are cheap and testing is automated. The thin band at the bottom is not the product but the parent stock for the next round, so imagine the picture repeating five or ten times.
What became clearer
WHAT CLEARED #When a target cannot be specified but can be recognised, recognition is enough. Directed evolution never solves the protein design problem; it sidesteps it by moving the chemist's knowledge out of the molecule and into the test, then letting cheap random variation and repeated local steps do the searching. The corollary is uncomfortable and exact: the quality of the result is bounded by the honesty of the filter, so a badly chosen assay does not fail loudly — it succeeds at the wrong thing.
Where to go next
ONWARD #- How a selection is engineered so that a cell's survival genuinely depends on the reaction of interest, rather than on a shortcut around it.
- Why promiscuous side activities exist at all, and what that implies about how new enzyme families arise.
Key terms
TERMS #| Term | What it means |
|---|---|
| Directed evolution | iterated rounds of random gene variation and selection or screening, used to obtain proteins with desired properties. |
| Error-prone PCR | a deliberately inaccurate gene-copying step used to generate a mutant library. |
| Genotype-phenotype linkage | keeping each protein physically associated with the gene that encoded it, so winners can be recovered and bred. |
| Promiscuous activity | a weak side reaction an enzyme catalyses incidentally, which is the usual starting signal for evolving a new function. |
| Fitness landscape | the mapping from sequence to performance whose ruggedness decides whether stepwise improvement can proceed. |
Every term the collection defines is gathered in the glossary.