VoiceRhythm

Shadowing

Reading aloud teaches you to pronounce words. Shadowing teaches you the music around them — where a fluent speaker speeds up, where they lean on a word, where they leave a gap. You borrow someone else's rhythm until it becomes yours.

Requires
Headphones
Methodology
shadowing-deterministic-v1
Score
Rhythm 40 · rate 30 · words 30
Lag
Reported, never penalised
The method

Why it works

Stress and intonation are almost impossible to describe and easy to copy. So you copy them. The narration plays in your headphones and you speak along with it, half a beat behind.

  • Prosody by imitation

    You are not told where the emphasis goes; you hear it and reproduce it. Imitation gets at the parts of delivery that instructions cannot reach.

  • Support removed in steps

    Audio, then the text, then the text while you speak, then nothing. Each stage drops one crutch instead of all of them at once.

  • A half-beat behind is correct

    Trailing the voice slightly is how shadowing is supposed to feel. Chasing perfect unison makes it worse — and the scorer knows this, which is why lag is measured and reported but never charged against you.

How a session runs

Step by step

  1. 1

    Listen to the narration once, then study the text.

  2. 2

    Practise along with it while the text is still visible.

  3. 3

    Take the final pass — headphones on, so only your voice is in the recording.

  4. 4

    Only that final take is measured; the listening and practice stages are never recorded and never scored.

Evaluation

How it's scored

This is the one methodology in the app whose inputs live on two clocks. The narration's timings come from a sidecar measured at normal speed; your take runs on real elapsed time. Rhythm and rate are derived from the audio envelope and those measured timings — no speech recogniser reports word-level timestamps on either platform, so none are used. The transcript enters only as a guard: without it, the optimal strategy would be to hum in time.

Deterministic · shadowing-sync

Rhythm match

40%

Your pauses matched against the narrator's, aligned on absolute time around an estimated constant lag rather than compared as a bare sequence of durations.

Rate match

30%

A tight, symmetric Gaussian around the pace the narration was heard at. Symmetric on purpose: you are hitting a rate someone else set, so running ahead is exactly as wrong as falling behind.

Word accuracy

30%

The Reading Aloud alignment, run against the passage — the guard that makes the rhythm score mean something.

The formula

Transcribed from the scorer that ships in the app. Every constant is pinned to this methodology version, so the same take always produces the same number.

Rhythm
referencePauses = narration gaps ≥ 250 ms
match window    = ±700 ms around the lag-adjusted position

hitRate = matched / reachedReferencePauses
quality = full credit at 0.6×–1.8× the reference length, else half

base    = 0.65 × hitRate + 0.35 × quality
rhythm  = base − extraPausePenalty        // proportional, capped at 25
Rate
targetWpm = narrationWords × 60000 / (narrationMs / playbackSpeed)
sigma     = 0.12 × targetWpm
rate      = 100 × exp(−(wpm − targetWpm)² / (2σ²))
Overall
overall = 0.40 × rhythm + 0.30 × rate + 0.30 × wordAccuracy

Details worth knowing

  • The extra-pause penalty is charged in proportion to how many pauses the narration itself holds, not as a flat fee per stray pause — otherwise a longer passage silently becomes a harder exercise, and a measured take that answered 24 of 31 pauses would score the same zero as one that never kept time at all.
  • A take that stopped mid-passage is graded against the part it actually shadowed. Pauses the narrator took after you stopped are not counted against anything.
  • Lag is a description of the take, not an error term. Negative lag means you ran ahead of the voice.

When no score is given

A weak guess is worse than no number, so the overall is withheld rather than estimated. What was measured stays visible either way.

  • No headphones → unscoreable. The narration would be in the recording alongside you and neither platform cancels it, so there is no separable evidence of what you actually said.
  • Fewer than 4 reference pauses reached, less than 15 seconds of active speech, or fewer than 10 recognised words → no overall.
  • No measured timing sidecar for the narration, a truncated speech timeline, or an unreliable transcription → no overall. The components you did earn stay visible; they just do not become a number that feeds a trend.

Reading the number

85–100Excellent
70–84Good
50–69Average
Below 50Needs practice
The coach review

What the AI adds

The review is written for shadowing specifically, and is warned about its characteristic errors:

  • How well you stayed with the narration — catching up, falling behind, or dropping words to keep up.

  • Whether you broke phrases where the narrator would have.

  • Whether you copied the narrator's emphasis and pitch movement or flattened it.

  • A consistent delay is explicitly not reported as hesitation — but trailing off at the ends of phrases while listening ahead is the real shadowing error, and that is worth telling you about.

Try Shadowing Yourself

Download Eloqo, open this exercise, and speak. That's the whole first session.

In closed testingGoogle PlayComing soon to theApp Store

Eloqo is in closed testing while we finish the public release.See the app in action →