Skip to main content
The voice stack determines how your agent listens, thinks, and speaks during a call. It controls how speech is processed, which model generates the response, and which voice engine turns that response into natural audio. Choosing the right stack matters because it affects latency, voice quality, reliability, multilingual performance, and how well the agent handles complex logic. RapidCall supports three voice stack options.

Gemini Studio

Gemini Studio is a complete text-to-speech stack from Google. It is a strong choice for English-language agents that need solid performance at a moderate cost. Typical latency is in the 700 to 1,200 ms range, and average cost is around $0.15 per minute. This option is best suited for use cases such as:
  • receptionist agents,
  • appointment booking,
  • simple outbound outreach,
  • and other structured call flows that do not depend on highly complex logic.

OpenAI Realtime

OpenAI Realtime is a speech-to-speech model built for very fast interactions. It offers very low latency, typically in the 150 to 500 ms range, and performs well across multiple languages. This makes it a strong fit for conversations where speed and responsiveness matter most. It is generally best for:
  • FAQ agents,
  • fast inbound conversations,
  • multilingual front-desk use cases,
  • and scenarios where natural responsiveness is a priority.
It is not the best choice for agents that need to follow complex multi-step logic or execute multiple actions with high reliability.

Custom Stack

The Custom Stack gives you full flexibility to choose your own language model and voice provider. This is the recommended option for more advanced use cases, especially when the agent needs to handle longer conversations, more complex prompts, or multiple actions during the call. It also supports 18+ languages and a wider range of voice and model combinations. Recommended combinations currently include:
  • Claude Sonnet + Cartesia for complex prompts, higher-quality reasoning, and multi-action workflows, at around $0.20 per minute,
  • Claude Haiku + Cartesia for more cost-efficient outbound workflows, typically around 15% cheaper while still keeping latency under 1,000 ms.

Which option should you choose?

As a general rule:
  • choose Gemini Studio for reliable, lower-complexity English agents,
  • choose OpenAI Realtime when speed is the main priority,
  • choose Custom Stack when you need the most control, flexibility, or workflow complexity.
If you are unsure which stack is right for your use case, the RapidCall team can help you choose and configure the best setup.