THIS EXPLANATION
THE ROOM
LAN·22 Language, Media & Communication 6 MIN · 7 STATIONS

Rare words in corpora

A Socratic walk-through of rare words in corpora — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does a dictionary need vastly more text to settle a rare word than a common one?

A dictionary editor wanting to know how the behaves has an embarrassment of evidence: any page of any book supplies dozens of instances. An editor wanting to settle nesh, or a chemical term, or a sense of cleave that may or may not still be live, can read a hundred million words of text and come away with three examples, none of them decisive.

The tempting reading is that this is just proportionality — rare words are rare, so of course they turn up less — which would make the problem merely tedious, solved by a corpus ten times bigger. It is not tedious in that way. Corpora have grown by orders of magnitude in living memory and the rare-word problem has not receded at all.

b

Reasoning it through

REASONING #

Start with the arithmetic, because it is unusually clean here and worth doing rather than recalling. Treat a word as occurring at some rate — a share of all running words. A word used once per million words of text will appear about a hundred times in a hundred-million-word corpus. A word used once per hundred million will appear about once. If an editor wants twenty independent examples of that second word, twenty divided by one-per-hundred-million gives two billion words. Nothing subtle has happened: the evidence about a word is the count of that word, and the corpus supplies it at a rate the editor does not control.

Notice what that kills. The relevant sample size for a word is not the size of the corpus; it is the number of occurrences the corpus happens to contain. This is where the ordinary sampling intuition — that a thousand respondents describe a nation as well as a town, because precision depends on how many you asked, not on how many there were — stops being useful. It does not stop being true: in a poll you choose n, whereas in a corpus you choose N and let the language decide n.

So far this is only expensive. Where does it become structural? Word frequencies are not a bell around some typical value: a handful of words take a large share of all running text and the remainder trails into an enormous tail of items used almost never. Because so much of the vocabulary sits out there, a corpus of any size contains a large stock of words seen exactly once — hapax legomena — and those are precisely the items no editor can settle, since one citation shows a word existed on one occasion and nothing else.

Now ask the decisive question: what happens to that stock when the corpus grows? If rare words were simply being flushed out by volume, the once-only stock should shrink toward nothing. It does not. Enlarging the corpus does two things at once — it converts some old hapaxes into well-attested words, and it admits a fresh crop of even rarer items that were absent before. The count of distinct word types keeps rising as text is added, with no sign of a ceiling, which is what you would expect from an open vocabulary constantly generating new compounds, names, derivations and errors. The frontier moves; it does not close. That is why a bigger corpus feels like it has not helped: it has helped with yesterday's rare words and handed you today's.

A second demand stacks on top, and it belongs to lexicography rather than statistics. An editor wants twenty independent occurrences — different authors, different years — because twenty hits from one obsessive writer establish an idiolect, not a word. Independence is scarcer than frequency, so the real requirement is harder than the arithmetic suggested.

I should mark the boundary of what I am claiming. The heavy-tailed shape of word frequencies and the persistence of once-only words are robust observations; the precise proportions — what share of types are hapaxes in a given corpus — vary substantially with the corpus, the language, and even the decision about what counts as one word, so I am declining to quote a figure that would look more solid than it is.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of dredging a river for coins. The common coins come up in every bucket, and after a few buckets you know their weight and mint precisely. A rare coin comes up once, and the only way to see a second is to dredge a great deal more river — during which you also bring up several coins you had never seen at all, each represented by a single specimen.

WHERE IT BREAKS DOWN

the coins in the river are a fixed stock being sampled, whereas a vocabulary is not fixed — speakers coin new words while the dredging is under way, so the supply of once-only items is being replenished from outside the corpus rather than merely being slow to surface.

d

Clarifying the model

THE MODEL #

Three refinements connect the pieces.

The first is that scarcity of examples, not of the word, is doing the work. Two words with the same overall rate can be unequal to settle: one confined to a single technical register is well attested in the right corpus and invisible in a general one. Corpus design can therefore beat corpus size, which is why lexicographers still run targeted reading programmes alongside billion-word databases.

The second separates two things that get run together. "Is this a word?" sounds like a question about the language. As dictionaries actually operate it is an editorial threshold: a policy about how many independent citations, across how many years, warrant an entry. That threshold is a defensible convention, not a discovery, and it can be tightened or loosened without anything about the language changing. The structural fact is the distribution of use; the prescriptive fact is where an editor draws the line on it. Confusing them produces the familiar and false claim that a word "is not really a word" until a dictionary admits it — which inverts the direction of evidence entirely.

The third is the falsification test, and it is worth stating sharply. If the difficulty really is a matter of rate, then the count of any given rare word should grow roughly in proportion to the corpus: multiply the text by ten and its attestations should multiply by about ten. That is checkable directly. What would refute the account is finding that a rare word's count grows much more slowly than the corpus, or not at all — which would mean the word is not being sampled at a stable rate and some other process is at work. And the open-vocabulary half has its own refuting observation: if the stock of once-only words fell steadily toward zero as corpora grew, the tail would be finite and the problem genuinely would be solved by volume.

e

A picture of it

THE PICTURE #
Rare words in corpora
Rare words in corpora Read across for how often the word occurs, and up for how much evidence the editorial decision needs -- two independent difficulties, which the picture exists to keep apart. The top-right is common but demanding: splitting a new sense off an old one needs many citations even for a word you meet daily. The left-hand column is where corpus size bites, and the top-left is the genuinely expensive cell -- rare and evidence-hungry -- which no realistic increase in general text volume reaches, and which is why targeted reading programmes still exist alongside enormous corpora. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/rare-words-in-corpora.md","sourceIndex":1,"sourceLine":4,"sourceHash":"53b2691369c175c013d7ab207a5a4ff3433945b50c6d3beb4a90ed028ad48a84","diagramType":"quadrantChart","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":720,"height":621},"qa":{"passed":true,"findings":[]}} Volume settles it Q1 Targeted reading Q2 Cheap to record Q3 Routine drafting Q4 Technical term in one journal One-off compound Regional dialect term New sense of a common verb Spelling of a common verb Rarely used Often used One citation suffices Many needed What makes a word hard to settle

How to readRead across for how often the word occurs, and up for how much evidence the editorial decision needs — two independent difficulties, which the picture exists to keep apart. The top-right is common but demanding: splitting a new sense off an old one needs many citations even for a word you meet daily. The left-hand column is where corpus size bites, and the top-left is the genuinely expensive cell — rare and evidence-hungry — which no realistic increase in general text volume reaches, and which is why targeted reading programmes still exist alongside enormous corpora.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The problem is not that rare words are rare; it is that the sample size for a word is the count of that word, so the editor cannot buy precision by buying text the way a pollster buys it by asking more people. And because the vocabulary is open, enlarging the corpus replenishes the tail as fast as it drains it. What remains is a design problem rather than a volume problem, plus an editorial threshold that ought not to be mistaken for a fact about the language.

g

Where to go next

ONWARD #
  • How Good-Turing estimation uses the count of once-only words to estimate the probability of words never seen at all.

Nearby on the shelf

4