Gemini Studio
Gemini Studio is a complete text-to-speech stack from Google. It is a strong choice for English-language agents that need solid performance at a moderate cost. Typical latency is in the 700 to 1,200 ms range, and average cost is around $0.15 per minute. This option is best suited for use cases such as:- receptionist agents,
- appointment booking,
- simple outbound outreach,
- and other structured call flows that do not depend on highly complex logic.
OpenAI Realtime
OpenAI Realtime is a speech-to-speech model built for very fast interactions. It offers very low latency, typically in the 150 to 500 ms range, and performs well across multiple languages. This makes it a strong fit for conversations where speed and responsiveness matter most. It is generally best for:- FAQ agents,
- fast inbound conversations,
- multilingual front-desk use cases,
- and scenarios where natural responsiveness is a priority.
Custom Stack
The Custom Stack gives you full flexibility to choose your own language model and voice provider. This is the recommended option for more advanced use cases, especially when the agent needs to handle longer conversations, more complex prompts, or multiple actions during the call. It also supports 18+ languages and a wider range of voice and model combinations. Recommended combinations currently include:- Claude Sonnet + Cartesia for complex prompts, higher-quality reasoning, and multi-action workflows, at around $0.20 per minute,
- Claude Haiku + Cartesia for more cost-efficient outbound workflows, typically around 15% cheaper while still keeping latency under 1,000 ms.
Which option should you choose?
As a general rule:- choose Gemini Studio for reliable, lower-complexity English agents,
- choose OpenAI Realtime when speed is the main priority,
- choose Custom Stack when you need the most control, flexibility, or workflow complexity.
