Launch Week 3: Five days of launches

Voice for AI Connections

Call your voice agent over phone, SIP, or WebRTC and evaluate it with simulated conversations.

Included on the Enterprise plan. Book a demo, opens in a new tab. Not included on the Team plan. Not included on the Starter plan. Not included on the Free plan.

Overview

When your AI app is a voice agent, you can configure your AI Connection to call it instead of sending it HTTP requests. During a multi-turn simulation, Confident AI places a real call to your agent, speaks each simulated user turn out loud, and transcribes what your agent says back. Every call is recorded, so you can listen to the conversation that produced each test case.

Confident AI supports three voice response modes:

  • Phone Call: dial a phone number and talk to your agent.
  • SIP Call: dial a SIP address and talk to your agent.
  • WebRTC: join your agent's WebRTC room and talk to it, either through your own signaling URL or through LiveKit.

You pick the mode in the AI App Endpoint section of your AI connection, the same place you'd pick HTTP Streaming or SSE Streaming.

Configuring Your Endpoint

What you enter as the endpoint depends on the response mode:

Response ModeEndpointExample
Phone CallA phone number with country code+1-415-555-0100
SIP CallA sip: or sips: addresssip:agent@example.com
WebRTCYour agent's signaling URLhttps://agent.example.com/offer
WebRTC (LiveKit)Your LiveKit URLwss://your-project.livekit.cloud

Since there's no request body or response to parse on a call, the Body, Output Parsing, and Resources tabs are hidden for voice connections. Your agent's actual output for each turn is the transcript of what it said.

Phone & SIP Calls

For Phone Call and SIP Call, enter the number or address your agent answers on. Confident AI places the call from its own line, so there are no credentials to configure.

  • Phone numbers must include the country code (e.g. +1-415-555-0100).
  • SIP addresses take the form sip:user@host, with an optional port and transport, for example sip:agent@example.com:5060;transport=tls. Supported transports are udp, tcp, and tls.

WebRTC

For WebRTC, Confident AI joins your agent as a participant and talks to it over audio. Choose how Confident AI should connect:

Enter your agent's signaling URL as the endpoint. It must start with https:// or wss://.

Any headers you add in the Headers tab are sent when Confident AI connects to your signaling URL, so this is where you'd put an API key or bearer token your agent expects.

Testing Your Connection

For voice connections, the Ping Endpoint button becomes Dial. Click Dial to place a test call to your agent. Confident AI connects, waits for your agent to greet the caller, and records what it hears.

Once the call ends you'll see:

  • The transcript of your agent's greeting
  • A recording of the greeting you can play back
  • How long the greeting and the call took

If the call connects but your agent doesn't say anything, you'll see Connected, but no greeting was heard. Voice simulations expect your agent to speak first, so make sure it greets the caller as soon as the call is answered.

✅ Done. Your voice connection is ready to run simulations.

Max Concurrent Calls

For voice connections, the Throttling tab shows Max Concurrent Calls instead of request concurrency: the maximum number of calls Confident AI places at the same time when simulating conversations with this AI connection.

Set this to what your agent and its telephony provider can handle. Each call ties up a line for the length of the conversation, so voice simulations take longer than text ones with the same number of goldens.

Personas

Personas shape how the simulated caller behaves. Navigate to Project Settings → Personas to create one. On top of a persona's Name, Characteristics, and Metadata, two settings only apply to voice calls.

Interruptions

Interruptions control how readily the persona talks over your agent:

  • Frequency: how often the persona interrupts. One of Off (default), Rare, Normal, or Frequent.
  • When both sides talk: what the persona does when it and your agent speak at the same time. One of Backs off, Adaptive (default), or Insists.

You can also set interruptions on an individual conversational golden. If the golden's persona sets its own interruptions, the persona's settings take precedence.

Background Noise

Upload a .wav or .mp3 file (up to 10 MB) under Upload background noise to play it underneath the persona's speech, for example street noise or a busy call center. Use the Volume slider (0 to 1, default 0.3) to control how loud it is relative to the persona's voice.

Voice Models

Navigate to Project Settings → Model Settings → Voice Models to choose the speech models the simulated caller uses.

  • Text-to-Speech (TTS): the model used to speak simulated user turns to your agent. Defaults to gpt-4o-mini-tts.
  • Speech-to-Text (STT): the model used to transcribe what your agent says.

By default, Confident AI simulates each user turn as a cascade: transcribe your agent with the STT model, generate the next user turn with the simulation model, then speak it with the TTS model.

Speech-to-Speech

Turn on Use real-time architecture under Speech-to-Speech (S2S) to replace that cascade with a single real-time caller. This gives you lower latency and more natural interruptions, closer to how a real caller behaves.

The real-time caller runs on fal. Pick Confident AI as the provider to use Confident AI's key, or pick fal and add your own fal API key under Model Settings → Credentials.

Running a Voice Simulation

There's nothing extra to set up at run time. Run a multi-turn evaluation on your dataset and select your voice connection. Each golden becomes one call.

While a test case is generating, its status shows where the call is:

  • Preparing caller...: the real-time caller is starting up (Speech-to-Speech only)
  • Dialing...: Confident AI is placing the call
  • In call...: the conversation is in progress

Your agent always speaks first. If a conversational golden's turns start with an assistant turn, Confident AI skips it and waits for your agent's greeting instead.

Once a test case finishes, open it to see:

  • Call Recording: the full recording of the call
  • Turn audio: each turn has its own audio clip alongside its transcript
  • Answered in: how long your agent took to respond to each turn
  • Interrupted: shown on turns where one side talked over the other

Next Steps

With your voice connection set up, run your first simulation.

Setting this up for your organization?Set up the controls your team needs before a wider rolloutTalk to us

Last updated on

Built byConfident AI