What drives latency
Total response latency is usually the combined result of three stages:
The exact number depends on the stack you are using, the complexity of the prompt, and the way the call settings are configured.
Common causes and fixes
Heavy language model
Larger models usually take longer to respond. If your use case is relatively structured - such as appointment booking, scripted outreach, or basic qualification - using a lighter model can reduce latency significantly without materially hurting performance. If latency is too high, one of the first things to test is switching from a heavier model to a lighter one.Long or unstructured prompt
Prompt size and prompt quality both affect inference speed. If the model has to process too much unnecessary text, duplicate instructions, or poorly structured guidance, response time increases. Remove anything that does not need to be in the prompt, move factual reference material into the Knowledge Base, and keep the prompt as clean and focused as possible.Response eagerness set too low
Response eagerness affects perceived latency even when the stack itself is reasonably fast. If response eagerness is set too low, the agent intentionally waits before speaking, which can make the call feel slower than it actually is. For most outbound and structured conversational use cases, higher eagerness settings usually create a tighter and more natural interaction.Voice provider speed
Different voice providers have different latency characteristics. The same agent can feel noticeably faster or slower depending purely on which text-to-speech provider is being used. If latency is important for the use case, compare multiple voice providers and test the difference directly in real calls.Stack choice
If the goal is the lowest possible latency, stack selection matters more than anything else. Some configurations prioritise speed, while others prioritise flexibility, reliability, or function execution. Faster options may feel more natural in conversation, but they can be less suitable for workflows that rely on more complex logic or multi-step actions.What to do first
If latency is consistently high, work through the changes in this order:- switch to a lighter model,
- simplify the prompt,
- increase response eagerness if appropriate,
- test a faster voice provider,
- compare the current stack against a lower-latency alternative.
Practical benchmark
As a general rule:- under 1,000 ms usually feels natural,
- 1,000-1,500 ms is noticeable but often acceptable,
- above 1,500 ms usually starts to feel slow in live conversation.
