Skip to main content
Testing is where the difference between a demo agent and a production agent becomes obvious. Most agents sound good on the first test call. What matters is how they perform when the conversation becomes less predictable — when the caller interrupts, goes silent, asks something unexpected, speaks unclearly, or reacts in a way the script did not anticipate. Do not skip testing. It is one of the highest-leverage parts of the build process.

The testing cycle

Every strong agent goes through the same cycle: build → test → fix → test again Most agents need two to five full iterations before they are ready for production. If your agent sounds perfect on the first try, you probably are not testing hard enough. A practical testing cycle usually looks like this:

1. Test the happy path first

Start with the ideal version of the conversation. The contact answers, follows the flow, responds clearly, and the agent completes the call exactly as intended. If the happy path does not work consistently, nothing else matters yet. Check that the agent:
  • follows the script in the correct order,
  • waits for responses before moving on,
  • uses the right dynamic variables,
  • triggers actions at the correct time,
  • and reaches the intended outcome cleanly.

2. Test objection paths

Next, test the objections your prompt is supposed to handle. If your agent has handlers for hesitation, pricing concerns, lack of urgency, requests for more information, or skepticism, trigger each one deliberately and verify that the agent:
  • acknowledges the objection properly,
  • redirects without arguing,
  • and returns to the main flow naturally.
This is where weak prompts often start to break down.

3. Test edge cases

After that, test the situations the script does not cover neatly. Examples include:
  • the contact says nothing for several seconds,
  • the contact gives an unclear answer,
  • the contact changes the subject,
  • the contact asks if the caller is AI,
  • the contact asks a question outside the prompt or knowledge base,
  • the contact gives incorrect information,
  • the contact asks for a transfer or callback unexpectedly.
A production-ready agent should not freeze, drift, or hallucinate in these cases. It should either handle them well or fall back cleanly.

4. Stress test the agent

Once the basic logic works, try to break it. Interrupt the agent mid-sentence.
Answer too quickly.
Mumble.
Ask the same question three times.
Change the subject.
Give contradictory information.
Speak over it.
Delay your response.
This is important because real contacts will not behave as neatly as test users do. A good agent must stay stable under messy, real-world conditions.

5. Make real phone calls

After testing inside the platform, call the agent from a real phone. This step matters because browser-based testing and real phone calls do not behave exactly the same way. In-browser tests are useful for logic and iteration, but real telephony adds compression, network variation, and audio conditions that can affect timing and speech recognition. Before declaring an agent ready, complete at least three to five real phone calls.

What to test

A proper test should cover four main areas.

Script adherence

Check whether the agent follows the conversation structure properly. Look for issues such as:
  • skipping steps,
  • combining steps that should be separate,
  • advancing too early,
  • repeating itself,
  • or ignoring the intended flow.
A strong prompt keeps the agent on track even when the caller is not perfectly cooperative.

Function execution

If the agent books appointments, sends SMS, transfers calls, updates records, or triggers webhooks, test every function path. Verify that:
  • the function fires at the right moment,
  • the correct data is passed,
  • the receiving system behaves as expected,
  • and the agent responds correctly to both success and failure.
Functions should be tested more rigorously than almost anything else. Poor wording is annoying. A broken function creates operational problems.

Voice quality and timing

Listen carefully to how the agent sounds, not just what it says. Check for:
  • unnatural pacing,
  • awkward delays,
  • interruptions at the wrong moment,
  • failure to yield when interrupted,
  • robotic tone,
  • or speech that feels too fast or too slow.
If the content is right but the delivery feels off, the problem is often in the call settings, not the prompt.

Edge-case handling

Test how the agent behaves when the conversation becomes less predictable. A few examples:

Using the Test Call tool

The built-in Test Call tool is the fastest way to iterate on an agent during setup. It is useful for testing:
  • prompt behaviour,
  • script flow,
  • objection handling,
  • dynamic variables,
  • and function execution.
Because it runs inside the platform, it is ideal for making quick adjustments and retesting immediately. What it does not fully simulate is real telephony. Browser audio is usually cleaner and lower-latency than a real phone line, so a call that sounds perfect in the test tool may still need tuning when tested over the phone network. Best practice:
  • use the Test Call tool for rapid iteration,
  • then validate the final version with real phone calls.

Reading test results

Every test call gives you data you can use to improve the agent.

Transcript

The transcript is the main debugging tool. Read it line by line and find the exact moment the agent drifted, responded badly, or missed the intended action. Then compare that moment to the relevant part of the prompt. That gap usually tells you what needs to change.

Recording

Listen to the full call, not just the text. Tone, pacing, awkward pauses, overlap, hesitation, and perceived naturalness are easier to catch in audio than in transcript form.

Latency

Latency tells you how responsive the agent feels. If latency is consistently too high, the call may feel unnatural even if the wording is good. In those cases, review the voice stack and settings rather than only rewriting the prompt.

Extracted data

If you are using post-call extraction, verify that the extracted fields match what actually happened in the conversation. If the transcript shows a booking but the extracted outcome says no booking, the extraction setup needs adjustment.

When is an agent ready for production?

An agent is ready for live calls when all of the following are true:
  • the happy path works consistently,
  • objection handlers behave as intended,
  • dynamic variables render correctly,
  • all functions work end to end,
  • real phone calls sound natural enough,
  • and edge cases are handled without freezing, drifting, or hallucinating.
One successful test call is not enough. The biggest mistake teams make is assuming an agent is ready because it performed well in one controlled conversation. Real callers will interrupt, hesitate, go off topic, speak unclearly, and behave in ways you did not expect. Testing has to reflect that.

Post-launch quality assurance

Testing does not stop once the agent goes live. The first 50 to 100 production calls usually reveal issues that earlier testing did not catch. Real callers surface real edge cases. After launch, review:
  • a sample of full call recordings,
  • transcripts from low-performing calls,
  • conversion rates across the intended funnel,
  • extracted field accuracy,
  • and sentiment or quality trends over time.
If the agent is supposed to book calls, review the calls where no booking happened. If sentiment drops, inspect the conversations behind it. If callers sound confused, adjust the prompt, handlers, or settings and test again. The best agents are not static. They improve continuously.

Best practice

A useful operating rhythm is:
  • test thoroughly before launch,
  • review the first production batch closely,
  • refine based on real conversations,
  • and continue improving over time.
Each iteration makes the agent more stable, more natural, and more effective.