VOICE SYSTEMS / INTERACTIVE WORKSHEET
Where does the pause go?
Budget the wait from the user’s last spoken sound to the first audible response. Start with a hypothetical serial path, then replace the inputs with measurements from your own traces.
SERIAL CRITICAL PATH / NOT A BENCHMARK
670 ms below your illustrative target. Largest stage: First usable text (400 ms).
- End-of-turn detection350 ms
- Remaining transcription180 ms
- Evidence retrieval120 ms
- First usable text400 ms
- First audio generation200 ms
- Delivery and playout80 ms
350 + 180 + 120 + 400 + 200 + 80 = 1,330 ms
Learn to replace assumptions with a trace ↗The boundary of the model
The sum includes only the serial durations you enter. Real systems may transcribe while a person speaks, retrieve speculatively, overlap text and speech, or introduce queueing and network delays. Model the actual critical path before interpreting a total. This worksheet does not predict a percentile, full-response duration, interruption responsiveness, lip synchronization or user satisfaction.
A lower total does not prove better turn-taking: overly aggressive end-of-turn detection can cut the user off. Compare traces with speech-repair outcomes and turn ownership, and inspect timeout budgets separately. Inputs are local to this page and are not stored or transmitted.
Original editorial model, prepared 19 September 2026. No live model, microphone or remote API is used.