THIS EXPLANATION
THE ROOM
BIO·27 Biology & Ecology 6 MIN · 8 STATIONS

Mark-recapture estimates

A Socratic walk-through of mark-recapture estimates — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

How can you count a population when you never see more than a handful of it?

Suppose you are asked how many trout live in a lake. You cannot drain it, and in a week of netting you handle sixty fish. Sixty is a fact about your net, not about the lake. So what could you possibly do with it?

Here is the move that makes the problem tractable. Stop trying to count animals and start trying to measure a ratio. You cannot see the whole population, but you can put a known number of marks into it and then ask the lake how diluted they came back.

b

Reasoning it through

REASONING #

Catch sixty trout, mark each one, put them back, and wait long enough for them to mix. Now net a second sample — say eighty fish — and count how many carry a mark. Suppose twelve do.

What does twelve tell you? It says that in the second catch, marked fish were 12 in 80, or 15 percent. If your second catch is a fair draw from the lake, then marked fish are about 15 percent of the lake too. And you know exactly how many marked fish there are: sixty. So sixty is 15 percent of the population, and the population is around four hundred.

Written out, that is the Lincoln-Petersen estimator: the first sample times the second, divided by the number of recaptures. Notice what it never required. It never required seeing most of the fish, or netting the same fraction twice, or knowing anything about the lake's volume. All it required was that the marks be a fair sample of the whole — which is a much smaller thing to assume than a full count.

Now press on the estimate, because this is inference under uncertainty and the estimate is not a number so much as a range. What if only two fish had come back marked instead of twelve? The arithmetic gives 2,400. But had you caught one fewer, it would have given 4,800, and one more, 1,600. The estimate hinges entirely on the recapture count, and when that count is small the answer swings wildly. This is why field biologists chase recaptures rather than raw catch: the precision lives in the overlap between the samples, not in their size. It is also why the plain formula is biased upward with few recaptures, and why practitioners use a corrected version — Chapman's — which adds one to each sample count and to the recaptures before dividing, precisely to tame the case where recaptures are near zero.

Then ask what would make the ratio lie, since every assumption is a place the estimate can fail. The population must be effectively closed during the study: births, deaths, and migration all change the denominator between samples. The marks must stay attached and stay readable. The two samples must be independent draws. And most treacherously, every animal must be equally catchable both times — which is exactly what a trapped animal often stops being. A raccoon that found food in a trap comes back for more, inflating recaptures and shrinking the estimate; a fish handled once may avoid nets, deflating recaptures and inflating the estimate. Trap-happy and trap-shy responses push the answer in opposite directions, and nothing inside the arithmetic reveals which one you have.

c

The analogy

THE ANALOGY #
THE FIGURE

Picture a jar of beans too full to count. Scoop out a cupful, paint every bean red, tip them back, shake the jar hard, and take a second cupful. If one bean in seven comes out red, the jar holds about seven times the number of beans you painted. The whole method is that one shake.

WHERE IT BREAKS DOWN

beans do not learn. Real animals remember the trap, drift to one end of the lake, lose their tags, breed, and die between the two scoops — so the shake you assumed happened is the assumption that actually fails, and the arithmetic cannot tell you it did.

d

Clarifying the model

THE MODEL #

Three refinements sharpen the picture. First, the estimate is not a count with a decimal point of doubt; it is a distribution, and it should always be reported with a confidence interval derived from the recapture count. A study reporting "412 trout" without a range has hidden the only thing worth arguing about.

Second, the closed-population assumption is a choice, not a fact of nature — and when it cannot hold, the answer is not to abandon the method but to extend it. Repeated sampling occasions let models such as Jolly-Seber estimate survival and recruitment alongside abundance, treating turnover as a quantity to measure rather than an error to avoid.

Third, unequal catchability is the assumption most worth designing against, and the standard defences are practical rather than mathematical: use two genuinely different capture methods, use passive marks the animal cannot learn to avoid, or use marks that require no capture at all — camera-trap coat patterns for tigers, photographed flukes for whales, DNA from hair snags or scat. The same logic runs well outside ecology. Public-health and census statisticians use overlapping incomplete lists the same way, treating the people who appear on two lists as the "recaptures" and estimating how many appear on none.

e

A picture of it

THE PICTURE #
Mark-recapture estimates
Mark-recapture estimates start at the parallelogram, the first catch, and follow the marks into the store node, the lake. The two diamonds are the assumptions the method actually rests on; take the "no" branch from the first and you land on the hazard node, where the arithmetic still returns a number that means nothing. The loop back from the second diamond is what field seasons are really spent doing. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/mark-recapture-estimates.md","sourceIndex":1,"sourceLine":4,"sourceHash":"e28ce8eb171902e2c54ff519c66d1ec99ef9c4c52ea0b293f5fc610cd101ad86","diagramType":"flowchart-v2","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":816,"height":1211},"qa":{"passed":true,"findings":[]}} yes no yes few Catch first sample Mark and release Marked fish now mixed in thelake Have the marks mixed evenly? Catch second sample Estimate biased, directionunknown Count how many carry a mark Enough recaptures to be stable? Population estimate with interval Widen interval or sample again
KINDSsourceprocessdecisionriskoutcomeconnector

How to readstart at the parallelogram, the first catch, and follow the marks into the store node, the lake. The two diamonds are the assumptions the method actually rests on; take the "no" branch from the first and you land on the hazard node, where the arithmetic still returns a number that means nothing. The loop back from the second diamond is what field seasons are really spent doing.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Counting was never the point. You measure a proportion you can see, anchor it to a number you set yourself, and let the two together imply the total you cannot see. What you buy with that trick is an estimate whose honesty lives entirely in its assumptions — which is why the recapture count, not the catch, is the number to watch.

g

Where to go next

ONWARD #
  • How Jolly-Seber and other open-population models turn births and deaths from a bias into a measurement.
  • Why non-invasive marks — photographs, DNA, acoustic signatures — have largely displaced physical tagging for rare species.
  • Multiple systems estimation: the same arithmetic applied to overlapping human records, and the sharp ethical questions that come with it.
h

Key terms

TERMS #
TermWhat it means
Lincoln-Petersen estimatorpopulation size estimated as the first sample times the second, divided by the number of marked animals recaptured.
Chapman estimatora bias-corrected form that adds one to each count before dividing, used when recaptures are few.
Closed populationone with no births, deaths, immigration, or emigration during the study period.
Trap-happy / trap-shya change in an animal's catchability caused by having been caught before, which biases the estimate down or up respectively.
Jolly-Seber modela family of models for open populations that estimates survival and recruitment along with abundance.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4