Skip to main content
Latency is the time between the contact finishing their sentence and the agent beginning its response. If the conversation feels slow, hesitant, or full of unnatural pauses, latency is usually the reason. RapidCall shows latency metrics for each call inside the Call Detail Panel, which makes it possible to diagnose the issue directly from live or test calls.

What drives latency

Total response latency is usually the combined result of three stages: The exact number depends on the stack you are using, the complexity of the prompt, and the way the call settings are configured.

Common causes and fixes

Heavy language model

Larger models usually take longer to respond. If your use case is relatively structured - such as appointment booking, scripted outreach, or basic qualification - using a lighter model can reduce latency significantly without materially hurting performance. If latency is too high, one of the first things to test is switching from a heavier model to a lighter one.

Long or unstructured prompt

Prompt size and prompt quality both affect inference speed. If the model has to process too much unnecessary text, duplicate instructions, or poorly structured guidance, response time increases. Remove anything that does not need to be in the prompt, move factual reference material into the Knowledge Base, and keep the prompt as clean and focused as possible.

Response eagerness set too low

Response eagerness affects perceived latency even when the stack itself is reasonably fast. If response eagerness is set too low, the agent intentionally waits before speaking, which can make the call feel slower than it actually is. For most outbound and structured conversational use cases, higher eagerness settings usually create a tighter and more natural interaction.

Voice provider speed

Different voice providers have different latency characteristics. The same agent can feel noticeably faster or slower depending purely on which text-to-speech provider is being used. If latency is important for the use case, compare multiple voice providers and test the difference directly in real calls.

Stack choice

If the goal is the lowest possible latency, stack selection matters more than anything else. Some configurations prioritise speed, while others prioritise flexibility, reliability, or function execution. Faster options may feel more natural in conversation, but they can be less suitable for workflows that rely on more complex logic or multi-step actions.

What to do first

If latency is consistently high, work through the changes in this order:
  1. switch to a lighter model,
  2. simplify the prompt,
  3. increase response eagerness if appropriate,
  4. test a faster voice provider,
  5. compare the current stack against a lower-latency alternative.

Practical benchmark

As a general rule:
  • under 1,000 ms usually feels natural,
  • 1,000-1,500 ms is noticeable but often acceptable,
  • above 1,500 ms usually starts to feel slow in live conversation.
If your recent calls are consistently above that range, the language model is usually the first place to optimise.