Shadowing
Reading aloud teaches you to pronounce words. Shadowing teaches you the music around them — where a fluent speaker speeds up, where they lean on a word, where they leave a gap. You borrow someone else's rhythm until it becomes yours.
- Requires
- Headphones
- Methodology
- shadowing-deterministic-v1
- Score
- Rhythm 40 · rate 30 · words 30
- Lag
- Reported, never penalised
Why it works
Stress and intonation are almost impossible to describe and easy to copy. So you copy them. The narration plays in your headphones and you speak along with it, half a beat behind.
Prosody by imitation
You are not told where the emphasis goes; you hear it and reproduce it. Imitation gets at the parts of delivery that instructions cannot reach.
Support removed in steps
Audio, then the text, then the text while you speak, then nothing. Each stage drops one crutch instead of all of them at once.
A half-beat behind is correct
Trailing the voice slightly is how shadowing is supposed to feel. Chasing perfect unison makes it worse — and the scorer knows this, which is why lag is measured and reported but never charged against you.
Step by step
- 1
Listen to the narration once, then study the text.
- 2
Practise along with it while the text is still visible.
- 3
Take the final pass — headphones on, so only your voice is in the recording.
- 4
Only that final take is measured; the listening and practice stages are never recorded and never scored.
How it's scored
This is the one methodology in the app whose inputs live on two clocks. The narration's timings come from a sidecar measured at normal speed; your take runs on real elapsed time. Rhythm and rate are derived from the audio envelope and those measured timings — no speech recogniser reports word-level timestamps on either platform, so none are used. The transcript enters only as a guard: without it, the optimal strategy would be to hum in time.
Deterministic · shadowing-sync
Rhythm match
40%Your pauses matched against the narrator's, aligned on absolute time around an estimated constant lag rather than compared as a bare sequence of durations.
Rate match
30%A tight, symmetric Gaussian around the pace the narration was heard at. Symmetric on purpose: you are hitting a rate someone else set, so running ahead is exactly as wrong as falling behind.
Word accuracy
30%The Reading Aloud alignment, run against the passage — the guard that makes the rhythm score mean something.
The formula
Transcribed from the scorer that ships in the app. Every constant is pinned to this methodology version, so the same take always produces the same number.
referencePauses = narration gaps ≥ 250 ms
match window = ±700 ms around the lag-adjusted position
hitRate = matched / reachedReferencePauses
quality = full credit at 0.6×–1.8× the reference length, else half
base = 0.65 × hitRate + 0.35 × quality
rhythm = base − extraPausePenalty // proportional, capped at 25targetWpm = narrationWords × 60000 / (narrationMs / playbackSpeed)
sigma = 0.12 × targetWpm
rate = 100 × exp(−(wpm − targetWpm)² / (2σ²))overall = 0.40 × rhythm + 0.30 × rate + 0.30 × wordAccuracyDetails worth knowing
- The extra-pause penalty is charged in proportion to how many pauses the narration itself holds, not as a flat fee per stray pause — otherwise a longer passage silently becomes a harder exercise, and a measured take that answered 24 of 31 pauses would score the same zero as one that never kept time at all.
- A take that stopped mid-passage is graded against the part it actually shadowed. Pauses the narrator took after you stopped are not counted against anything.
- Lag is a description of the take, not an error term. Negative lag means you ran ahead of the voice.
When no score is given
A weak guess is worse than no number, so the overall is withheld rather than estimated. What was measured stays visible either way.
- No headphones → unscoreable. The narration would be in the recording alongside you and neither platform cancels it, so there is no separable evidence of what you actually said.
- Fewer than 4 reference pauses reached, less than 15 seconds of active speech, or fewer than 10 recognised words → no overall.
- No measured timing sidecar for the narration, a truncated speech timeline, or an unreliable transcription → no overall. The components you did earn stay visible; they just do not become a number that feeds a trend.
Reading the number
What the AI adds
The review is written for shadowing specifically, and is warned about its characteristic errors:
How well you stayed with the narration — catching up, falling behind, or dropping words to keep up.
Whether you broke phrases where the narrator would have.
Whether you copied the narrator's emphasis and pitch movement or flattened it.
A consistent delay is explicitly not reported as hesitation — but trailing off at the ends of phrases while listening ahead is the real shadowing error, and that is worth telling you about.
Try Shadowing Yourself
Download Eloqo, open this exercise, and speak. That's the whole first session.
Eloqo is in closed testing while we finish the public release.See the app in action →