← All practical tools

VOICE SYSTEMS / INTERACTIVE WORKSHEET

Where does the pause go?

Budget the wait from the user’s last spoken sound to the first audible response. Start with a hypothetical serial path, then replace the inputs with measurements from your own traces.

Six stages. One explicit assumption.

All stages below run in sequence in this model. Do not enter overlapping durations twice. Every default is illustrative—not a recommended target or measured product result.

SERIAL CRITICAL PATH / NOT A BENCHMARK

1,330milliseconds to first audio

670 ms below your illustrative target. Largest stage: First usable text (400 ms).

  1. End-of-turn detection350 ms
  2. Remaining transcription180 ms
  3. Evidence retrieval120 ms
  4. First usable text400 ms
  5. First audio generation200 ms
  6. Delivery and playout80 ms

350 + 180 + 120 + 400 + 200 + 80 = 1,330 ms

Learn to replace assumptions with a trace ↗

The boundary of the model

The sum includes only the serial durations you enter. Real systems may transcribe while a person speaks, retrieve speculatively, overlap text and speech, or introduce queueing and network delays. Model the actual critical path before interpreting a total. This worksheet does not predict a percentile, full-response duration, interruption responsiveness, lip synchronization or user satisfaction.

A lower total does not prove better turn-taking: overly aggressive end-of-turn detection can cut the user off. Compare traces with speech-repair outcomes and turn ownership, and inspect timeout budgets separately. Inputs are local to this page and are not stored or transmitted.

Original editorial model, prepared 19 September 2026. No live model, microphone or remote API is used.

Find your next good decision.

Start typing to explore the guides.

76 sourced guides · Escape to close