Delayed auditory feedback
A Socratic walk-through of delayed auditory feedback — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does hearing your own voice a fraction of a second late stop you speaking?
Put on headphones, speak into a microphone, and have your own voice played back to you about a fifth of a second late. Within a sentence or two most people slow to a crawl, repeat syllables, stretch vowels, and get louder. Bernard Lee described this in 1950 and called it an artificial stutter.
Here is what should bother us. Nothing was taken away. Your tongue works, your lungs work, you still know the sentence. All that changed is the timing of information you already had. So before we ask what a delay does, let us ask something smaller: why should speaking depend on hearing yourself at all?
Reasoning it through
REASONING #Start with the obvious test. Cover your ears, or stand in a very noisy room, and speak. You can still do it. So hearing is not the thing that drives the muscles. Something else is issuing the commands.
Then ask how fast speech actually is. Syllables come at roughly four to seven a second, and the individual articulatory gestures inside them are shorter still. Now suppose the loop really did run through your ears: sound leaves your mouth, reaches your ear, gets analysed, and a correction comes back down. Would that be quick enough to steer a gesture that is already over? Almost certainly not — the round trip is on the order of the gesture itself.
So the system cannot be waiting on what it hears. The mainstream account is that it runs ahead: the motor system issues a command and, at the same time, predicts what that command should sound like. Hearing is then used to check the prediction, not to generate the movement. A mismatch means something drifted — a cold, a mouthful of food, a room that swallows your voice — and the system nudges rate, loudness, or articulation to bring things back.
Notice what that makes the ear: not a driver but a comparator. And a comparator is only meaningful if the two things it compares belong to the same moment.
Now put the delay back in. At 200 milliseconds — roughly the length of a syllable, and about where disruption peaks in the classic studies — what arrives at your ear is a syllable you have already finished, and it lands while the next one is being produced. The predictor says one thing; the ear reports something else entirely. There is a large error, and it is real: the signals genuinely disagree.
What does the loop then do? It corrects. But which syllable does the correction land on? Not the one that generated the error — that one is gone. It lands on whatever you are saying now. And that new syllable, corrected on the basis of a stale complaint, comes back late as well, producing another error. Slow down, get louder, repeat the syllable, and every one of those responses feeds the mismatch again.
That is the whole mechanism, and it is the classic recipe for an unstable control loop: a real error signal, a nonzero correction gain, and a loop delay comparable to the timescale of the thing being controlled. The loop does not fail because it is weak. It fails because it is fast and confident about information that is out of date.
Does that predict anything we can check? It predicts the disruption should be worst near a syllable's length, not simply worse the longer the delay. That is what is observed: very short delays are harmless, and delays beyond roughly half a second are heard as a separate echo, something you can dismiss as another voice rather than as your own mouth misbehaving. And it predicts that masking your own voice with noise should help, because a loop with no signal at least issues no false corrections. It does.
One honesty note: exactly which computation breaks is still argued. The prediction-and-comparator story above is the standard one, but some researchers put more weight on a mismatching stream of speech simply capturing attention. And there is a genuine puzzle sitting beside it — for many people who stutter, delayed or otherwise altered feedback makes speech more fluent, which is why such devices exist. Why the same manipulation should break fluent speakers and help stuttering ones is not settled.
The analogy
THE ANALOGY #Think of a shower whose hot water takes ten seconds to arrive. You are too cold, so you turn the tap. Nothing happens, so you turn it further. Then the first correction lands, scalding, so you swing it back hard — and now the second correction is on its way behind it. You oscillate, not because you are bad at judging temperature, but because you are judging a temperature that left the boiler before your last two adjustments existed.
you can beat the shower by deliberately waiting between adjustments, but the speech loop is not under that kind of voluntary control — and the shower simply gets worse the longer the lag, whereas speech is disrupted most at a delay near the length of a syllable and recovers when the lag grows long enough to hear as an echo.
Clarifying the model
THE MODEL #The tempting summary is "your brain gets confused by the echo". That is close, but it misplaces the problem. The system is not confused about what it hears; it hears your voice perfectly well. What it cannot do is stamp it with the right time.
The second correction worth making is about direction. It is easy to assume feedback drives speech and that damaging it removes the steering. But silence — masking noise, or a soundproof room — barely hurts, while a mis-timed signal is devastating. A wrong loop is worse than no loop. That is a general property of control systems and not a quirk of talking.
And this is why the effect is a laboratory tool rather than a curiosity. If a fifth of a second can reduce a fluent adult to syllable repetition, then whatever is being disturbed is doing real work in ordinary speech, continuously, below awareness.
A picture of it
THE PICTURE #How to readstart at the data node at the top and follow the two arrows out of the motor-command box — one goes to the internal prediction, the other out through your mouth. They meet at the diamond. Take the right-hand branch when the delayed report disagrees, and follow the back-edge from the shaded correction node: it returns to the motor commands, which is the loop that never settles.
What became clearer
WHAT CLEARED #Speaking is not steered by hearing, but it is audited by hearing, and the audit only works if the report and the prediction refer to the same instant. Delay the report by about one syllable and you have not removed a sense — you have handed a live control loop a confident, precisely wrong error signal, and watched it chase it.
Where to go next
ONWARD #- The Lombard effect: why we automatically raise our voice in noise, and what that says about the same loop.
- Why altered feedback often improves fluency for people who stutter.
- Delayed visual feedback in typing and drawing — does the same instability appear?
- Loop delay and gain in engineered control systems, where the instability condition is stated formally.
Key terms
TERMS #| Term | What it means |
|---|---|
| Delayed auditory feedback (DAF) | hearing one's own voice played back after a short lag, typically 50-500 ms. |
| Forward model | an internal prediction of the sensory consequences of a motor command, used for comparison against what actually arrives. |
| Loop delay | the time for a correction to travel round a feedback loop; instability follows when it approaches the timescale being controlled. |
| Lombard effect | the involuntary increase in vocal effort when background noise rises. |
Every term the collection defines is gathered in the glossary.