THIS EXPLANATION
THE ROOM
LAN·03 Language, Media & Communication 6 MIN · 8 STATIONS

Categorical perception of speech sounds

A Socratic walk-through of categorical perception of speech sounds — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does a smoothly changing sound flip abruptly from one consonant to another?

Build a set of synthetic syllables that differ in one continuous quantity: the delay between releasing the lips and the vocal folds starting to buzz. Set that delay to zero and English listeners hear ba. Stretch it to sixty milliseconds and they hear pa. Now step through in five-millisecond increments and ask what they hear at each step.

You might expect the reports to slide — more and more p-ish, some intermediate thing in the middle. They do not. Listeners report ba at every step until somewhere near the high twenties, and then pa at every step after. The stimulus changed by a smooth ramp; the percept changed by a cliff. Something between ear and report is converting a gradient into a decision.

b

Reasoning it through

REASONING #

First, be clear about the physical quantity. Voice onset time is a genuine continuum with no natural joint in it: the gap between the burst of a stop release and the onset of vocal fold vibration can take any value. Nothing in acoustics privileges 25 ms over 24. So the cliff cannot be in the signal.

Second, ask what the listener's task actually is. It is not to measure the delay. It is to identify which word was said — bat or pat — and there is no third option that lies between them. The output of the system has to be one of a small set of discrete symbols, because that is what the rest of the language machinery consumes. So somewhere between a continuous input and a discrete output, a threshold has to be applied. The only question was where, and how sharply.

Third, ask what such a threshold predicts if it really exists, because the identification curve alone is weak evidence — you could get a sharp curve from a forced choice even if perception were fully gradient. The interesting prediction is about discrimination. Take two stimuli that differ by exactly 20 ms and ask a listener whether they are the same or different. If the listener is measuring the physical delay, the pair should be equally discriminable wherever you place it on the continuum. If the listener is applying a threshold and then hearing the label, pairs straddling the boundary should be easy and pairs sitting inside one category should be near-impossible.

That is the classic result: discrimination peaks at the identification boundary and drops to near chance within categories. Lisker and Abramson's voice onset time work in the 1960s and 1970s established it for stops, following the Haskins Laboratories findings on synthetic speech in the late 1950s, and the pattern — sharp identification plus boundary-peaked discrimination — is what "categorical perception" names.

Now push on it, because the story as told is too clean. Where is the boundary? Not at a universal constant. English puts the ba-pa boundary around 25 to 30 ms; Spanish, whose voiced stops are prevoiced, puts it far lower; Thai contrasts three categories along the same dimension and needs two boundaries. So the threshold is learned from the language's own distribution of sounds, and infants' early sensitivity narrows during the first year to match the contrasts their language uses.

And the boundary moves within a listener too. Speak faster and it shifts, because the same delay means something different at a different rate. Lexical context shifts it: Ganong showed that an ambiguous sound is heard as whichever consonant makes a real word. Finally, the strong claim that within-category differences are simply invisible has not survived — reaction times, eye movements to pictures, and priming all show listeners retain graded sensitivity inside a category. The honest current statement is that the mapping is steep and category-biased rather than truly discrete, and that much of the apparent all-or-nothing quality arises at the decision stage rather than in the auditory system.

c

The analogy

THE ANALOGY #
THE FIGURE
Think of a thermostat wired to a heater. The room temperature is a smooth continuum, and the sensor tracks it faithfully, but the heater has only two states. As the temperature drifts down past the set point, the output does not fade gradually — it flips. Watch only the heater and you would conclude the room's temperature comes in two values. The gradient is real; it just meets a device whose job is to produce a decision.
WHERE IT BREAKS DOWN

the thermostat truly discards everything but the on-off decision, whereas listeners demonstrably retain graded information about where within a category a sound fell — the flip is a strong bias in the read-out, not a genuine erasure of the underlying continuum.

d

Clarifying the model

THE MODEL #

The misconception to head off is that categorical perception shows speech sounds are physically discrete. They are not. What is discrete is the phonological inventory a language uses, and perception has been tuned to deliver that inventory reliably from a noisy, variable signal in which no two productions of b are alike. Placing a steep threshold where the language's two distributions cross is close to the optimal way to recover the intended category, so the cliff is a sensible solution to a classification problem rather than a quirk of the ear.

It also explains why categorical perception is strongest for stop consonants and much weaker for steady-state vowels, which listeners discriminate within category quite well. Vowels are long, spectrally rich and speaker-variable; stops are brief transients whose cues are fleeting. The more the identity of a sound has to be committed quickly from thin evidence, the more the system behaves like a switch.

e

A picture of it

THE PICTURE #
Categorical perception of speech sounds
Categorical perception of speech sounds the two states are what the listener reports, and voice onset time is the continuous quantity driving the transition between them; there is no intermediate state, so a smooth increase in the delay produces no gradual percept, only a crossing -- and the note records that the crossing point is set by the listener's language and nudged by context rather than fixed by acoustics. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/categorical-perception-of-speech-sounds.md","sourceIndex":1,"sourceLine":4,"sourceHash":"3b2fc470e766b26e4a03d243011527de1b08abc49c17662b613e9fae5d201f4a","diagramType":"stateDiagram","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":720,"height":624},"qa":{"passed":true,"findings":[]}} voice onset delay growspast about 30 ms delay shrinks back belowthe boundary Heard as the b sound Heard as the p sound The boundary is learned,not fixed. Speech rate andreal words both shift it.
KINDSconnectorfeedback loop

How to readthe two states are what the listener reports, and voice onset time is the continuous quantity driving the transition between them; there is no intermediate state, so a smooth increase in the delay produces no gradual percept, only a crossing — and the note records that the crossing point is set by the listener's language and nudged by context rather than fixed by acoustics.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Speech perception has to hand a discrete symbol to the rest of the language system, so a continuous acoustic cue meets a learned threshold placed where the language's own categories divide — and a threshold on a smooth ramp produces exactly what the experiments show: flat identification either side, a cliff at the crossing, and discrimination that is sharp there and poor within.

g

Where to go next

ONWARD #
  • How infants' category boundaries narrow during the first year, and what happens to contrasts their language does not use.
  • The Ganong effect and other top-down shifts: how far lexical knowledge can move a phonetic boundary.
  • Whether categorical perception is special to speech, given that similar effects appear for musical intervals, colours and familiar faces.
h

Key terms

TERMS #
TermWhat it means
Voice onset time (VOT)the interval between the release of a stop consonant and the start of vocal fold vibration.
Categorical perceptionthe pattern of sharp identification plus discrimination that peaks at the category boundary and is poor within categories.
Phoneme boundarythe value of a continuous cue at which listeners switch from one category label to the other.
Prevoicingvocal fold vibration beginning before the stop release, giving a negative VOT, as in Spanish voiced stops.
Ganong effectthe tendency to hear an ambiguous speech sound as whichever category yields a real word.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4