RedFlag

Cognitive load in speech: pauses, fillers and speaking rate

← All articles · By the RedFlag team · September 4, 2026

Speech is produced in real time by a brain that is also doing other things — retrieving memories, monitoring the listener, planning the next clause. When those demands spike, the production system shows it. This is the cognitive-load account of speech, and it is the sturdiest scientific footing this whole field has.

Pauses are where thinking lives

Frieda Goldman-Eisler's foundational work in the 1950s–60s established that hesitation pauses are not noise: they cluster exactly where the cognitive work is — before unpredictable words, at planning boundaries, ahead of complex clauses. Fluent speech runs on pre-planned stretches; pausing is the sound of planning happening live. Later work on speech under load confirmed the package: under strain, speakers slow down, pause more (and differently), and their pitch range narrows.

Fillers are a load gauge, not a character flaw

"Um" and "uh" mark planning trouble, and they respond to load: harder questions and less rehearsed answers draw more of them. Clark and Fox Tree (2002) argued they function almost as words — signals to the listener that a delay is coming. The informative pattern is never a count in isolation; it is a change — a normally fluent speaker suddenly hedging every clause, or, just as telling, a normally hesitant speaker turning suspiciously fluent on the one answer that should require thought (the rehearsal signature).

The deception connection runs through load

Vrij, Fisher and Blank (2017) meta-analyzed the "cognitive approach": deceiving is often more demanding than truth-telling (invent, maintain consistency, suppress the truth, monitor the listener), and interventions that raise load amplify the observable differences between truthful and deceptive speakers. Two honest caveats: some lies are rehearsed and therefore cheap to produce, and plenty of truthful moments are demanding — difficult recall, emotional topics, speaking in a second language. Load signals mark effort, never lies.

A surprising continuity effect

One pattern in RedFlag's own labeled reference data cuts against intuition: flagged stretches are often more continuously voiced than the speaker's own norm — under pressure, some speakers do not fall silent but monologue through the tough spot, filling the space where a pause would invite a follow-up question. Direction held on 10 of 14 labeled videos. It earned a place in the verdict as a person-relative signal: voiced fraction against the video's own median.

What a careful listener can do

That is also exactly the design brief behind RedFlag: it measures speaking rate, filler ratio, pauses, pitch and blink signals against the speaker's own baseline and flags divergent moments in the video timeline. RedFlag is not a lie detector — the method and its limits are public at how RedFlag scores a video.

References

See the signals for yourself. Paste any public YouTube video with a single speaker into the RedFlag analyzer — the face and voice analysis runs in your browser, free, no signup. RedFlag is not a lie detector; it shows you where the signals diverge and leaves the judgment to you.

Open the analyzer