The error threshold on genome length
A Socratic walk-through of The error threshold on genome length — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why can a genome never grow longer than its own copying accuracy permits?
Genomes vary enormously in length, and the variation is not random. RNA viruses top out around thirty thousand bases. Bacteria run to millions. We run to billions. If a longer genome buys more capability, why did the RNA viruses not simply get longer?
The tempting answer is that they had no need, or no time. I would like to test a harder possibility instead: that there is a ceiling, that the ceiling is set by nothing more than how accurately the thing copies itself, and that no amount of selection can raise it.
Reasoning it through
REASONING #Let us treat a genome as a message being sent through a noisy channel, where the transmission is replication and the noise is mutation. Ask a simple question: how many errors does one copy introduce?
If the per-base error rate is u and the genome is L bases long, the expected number of mistakes per copy is simply u times L. So the interesting quantity is never the error rate alone and never the length alone — it is their product, the mutations per genome per replication.
Now what happens as that product approaches one? Roughly, every offspring differs from its parent somewhere. The exact sequence that natural selection has painstakingly assembled — the best one, the master sequence — is copied faithfully only rarely, and it is being produced anew from its own mutants even more rarely.
Here is where you have to be careful, because there is an obvious objection: surely selection just keeps killing the bad copies? True, and selection is doing exactly that. But selection can only act on what it can distinguish, and it can only preserve the master sequence if that sequence is regenerated faster than it is degraded. Manfred Eigen worked this out formally in 1971. The population settles not on one sequence but on a cloud of related mutants — a quasispecies — and the master sequence persists in that cloud only while u times L stays below roughly the natural logarithm of its selective advantage over the average mutant.
That logarithm is the reason the answer is so blunt. A tenfold fitness advantage gives you a budget of about 2.3; a hundredfold gives about 4.6. Because the logarithm crushes even large fitness differences into small numbers, the practical constraint is that u times L must be of order one, so L can be no more than roughly one over u. Copy at one error in ten thousand bases and you may carry about ten thousand bases. No more, whatever the selection pressure.
And past that point, what happens? Not gradual decay. The transition is sharp: heritable information dissolves into the mutant cloud, and the population drifts through sequence space rather than sitting on a peak. Eigen called this the error catastrophe.
Now test the prediction against the world, because this is where it gets satisfying. RNA viruses copy at roughly one error in ten thousand to a hundred thousand bases and carry genomes of about ten thousand bases — sitting right on the line. DNA organisms with proofreading polymerases and mismatch repair achieve error rates orders of magnitude lower, and their genomes are correspondingly longer. And the pointed case: coronaviruses have the largest known RNA genomes, near thirty thousand bases, and they are the RNA viruses that carry a proofreading exonuclease, nsp14-ExoN. Disable it experimentally and mutation rates rise while the genome becomes unstable. Does the correlation between fidelity machinery and permitted length look like a coincidence to you?
But notice the trap this sets. To copy more accurately you need proofreading enzymes. Proofreading enzymes must be encoded, which requires a longer genome. A longer genome requires better fidelity. Eigen named this circularity his paradox, and how the first replicators escaped it — through compartmentalisation, cooperating short sequences, or something not yet proposed — is a genuinely open question in origin-of-life research, not a solved one.
The analogy
THE ANALOGY #Imagine copying a book by hand, with each copy made from the previous one and no original to check against. If you make one error every ten thousand characters, a thousand-character pamphlet survives many generations essentially intact — most copies are perfect. A ten-thousand-character chapter accumulates roughly one error per generation, and errors are never removed. A million-character novel is unreadable within a few copies. There is a length above which a manuscript simply cannot be transmitted at that standard of scribe, and no amount of enthusiasm for long novels changes it.
the scribe has no editor, whereas a population does — selection culls the worst copies each generation, which is exactly why the real limit is not "zero errors" but a budget of order one mutation per genome, and why the threshold depends on how much better the master sequence is than its rivals.
Clarifying the model
THE MODEL #Three qualifications, since this is a model that is often overstated.
First, the threshold is not a hard wall in nature. The clean version assumes an unchanging fitness landscape, a well-mixed asexual population and one distinct master sequence. Recombination, sex, population structure and neutral mutations all shift where the edge falls. The scaling L of order one over u is robust; the precise coefficient is not.
Second, "error catastrophe" is a contested label for what happens to a real virus pushed over the edge. Mutagenic drugs such as ribavirin and favipiravir do drive viral populations to extinction, but most workers now describe the effect as lethal mutagenesis — simply too many lethal mutations for the population to sustain itself — rather than as Eigen's information-loss transition. The two are distinct mechanisms with similar outcomes, and the distinction is still argued.
Third, sitting near the threshold is not obviously a failure. A high mutation rate is also how an RNA virus explores antigenic space fast enough to outrun immunity, so the observed rates may reflect a fidelity-adaptability balance rather than a limitation the virus would remove if it could.
A picture of it
THE PICTURE #How to readEach bar is the genome length as a power of ten, and the line above it is one divided by the per-base error rate — the length that fidelity alone would permit. The vertical gap between line and bar is the safety margin. Read left to right and watch that gap open: it is essentially zero for a plain RNA virus, which is why that group cannot grow longer, and it widens only where better copying machinery appears.
What became clearer
WHAT CLEARED #Genome length is not a free parameter that organisms grow into as they need it. It is capped by an information constraint — mutations per genome per copy must stay of order one, or heritable information dissolves faster than selection can rebuild it. Every large genome in the world is therefore evidence of an earlier investment in copying accuracy, and the ceiling on RNA viruses is not a lack of ambition but a limit on what their polymerases can hold.
Where to go next
ONWARD #- How compartmentalisation and cooperating replicators are proposed to break Eigen's paradox.
- Why Drake's rule — a roughly constant mutation rate per genome per replication across microbes — falls out of the same argument.
- How the fidelity-versus-adaptability trade-off shapes drug resistance.
Key terms
TERMS #| Term | What it means |
|---|---|
| Error threshold | the maximum product of mutation rate and genome length above which the fittest sequence can no longer be maintained by selection. |
| Quasispecies | the cloud of closely related mutant sequences a replicating population actually consists of, rather than a single genotype. |
| Error catastrophe | the loss of heritable information that follows when a population crosses the error threshold. |
| Lethal mutagenesis | extinction caused by a drug-induced excess of lethal mutations, mechanistically distinct from information loss. |
| Proofreading exonuclease | an enzyme activity that excises a misincorporated base, raising replication fidelity by orders of magnitude. |
| Eigen's paradox | accurate copying needs enzymes, enzymes need a long genome, and a long genome needs accurate copying. |
Every term the collection defines is gathered in the glossary.