Voice for AI Connections
Call your voice agent over phone, SIP, or WebRTC and evaluate it with simulated conversations.
Overview
When your AI app is a voice agent, you can configure your AI Connection to call it instead of sending it HTTP requests. During a multi-turn simulation, Confident AI places a real call to your agent, speaks each simulated user turn out loud, and transcribes what your agent says back. Every call is recorded, so you can listen to the conversation that produced each test case.
Confident AI supports three voice response modes:
- Phone Call: dial a phone number and talk to your agent.
- SIP Call: dial a SIP address and talk to your agent.
- WebRTC: join your agent's WebRTC room and talk to it, either through your own signaling URL or through LiveKit.
You pick the mode in the AI App Endpoint section of your AI connection, the same place you'd pick HTTP Streaming or SSE Streaming.
Configuring Your Endpoint
What you enter as the endpoint depends on the response mode:
| Response Mode | Endpoint | Example |
|---|---|---|
| Phone Call | A phone number with country code | +1-415-555-0100 |
| SIP Call | A sip: or sips: address | sip:agent@example.com |
| WebRTC | Your agent's signaling URL | https://agent.example.com/offer |
| WebRTC (LiveKit) | Your LiveKit URL | wss://your-project.livekit.cloud |
Since there's no request body or response to parse on a call, the Body, Output Parsing, and Resources tabs are hidden for voice connections. Your agent's actual output for each turn is the transcript of what it said.
Phone & SIP Calls
For Phone Call and SIP Call, enter the number or address your agent answers on. Confident AI places the call from its own line, so there are no credentials to configure.
- Phone numbers must include the country code (e.g.
+1-415-555-0100). - SIP addresses take the form
sip:user@host, with an optional port and transport, for examplesip:agent@example.com:5060;transport=tls. Supported transports areudp,tcp, andtls.
WebRTC
For WebRTC, Confident AI joins your agent as a participant and talks to it over audio. Choose how Confident AI should connect:
Enter your agent's signaling URL as the endpoint. It must start with https:// or wss://.
Any headers you add in the Headers tab are sent when Confident AI connects to your signaling URL, so this is where you'd put an API key or bearer token your agent expects.
If your agent runs on LiveKit, Confident AI can join a room in your LiveKit project directly:
- Enter your LiveKit URL as the endpoint (e.g.
wss://your-project.livekit.cloud) - Open the Authentication tab and set Authentication Type to LiveKit
- Fill in your LiveKit credentials:
| Field | Description |
|---|---|
| API Key | The API key of the LiveKit project your agent runs in. Found under Settings → Keys in LiveKit Cloud. |
| API Secret | The API secret paired with the key. Used to mint the token the simulated caller joins the room with. |
| Agent Name (optional) | Only for agents that use explicit dispatch. Leave empty when your agent joins every new room automatically. |
Testing Your Connection
For voice connections, the Ping Endpoint button becomes Dial. Click Dial to place a test call to your agent. Confident AI connects, waits for your agent to greet the caller, and records what it hears.
Once the call ends you'll see:
- The transcript of your agent's greeting
- A recording of the greeting you can play back
- How long the greeting and the call took
If the call connects but your agent doesn't say anything, you'll see Connected, but no greeting was heard. Voice simulations expect your agent to speak first, so make sure it greets the caller as soon as the call is answered.
✅ Done. Your voice connection is ready to run simulations.
Max Concurrent Calls
For voice connections, the Throttling tab shows Max Concurrent Calls instead of request concurrency: the maximum number of calls Confident AI places at the same time when simulating conversations with this AI connection.
Set this to what your agent and its telephony provider can handle. Each call ties up a line for the length of the conversation, so voice simulations take longer than text ones with the same number of goldens.
Personas
Personas shape how the simulated caller behaves. Navigate to Project Settings → Personas to create one. On top of a persona's Name, Characteristics, and Metadata, two settings only apply to voice calls.
Interruptions
Interruptions control how readily the persona talks over your agent:
- Frequency: how often the persona interrupts. One of Off (default), Rare, Normal, or Frequent.
- When both sides talk: what the persona does when it and your agent speak at the same time. One of Backs off, Adaptive (default), or Insists.
You can also set interruptions on an individual conversational golden. If the golden's persona sets its own interruptions, the persona's settings take precedence.
Background Noise
Upload a .wav or .mp3 file (up to 10 MB) under Upload background noise to play it underneath the persona's speech, for example street noise or a busy call center. Use the Volume slider (0 to 1, default 0.3) to control how loud it is relative to the persona's voice.
Voice Models
Navigate to Project Settings → Model Settings → Voice Models to choose the speech models the simulated caller uses.
- Text-to-Speech (TTS): the model used to speak simulated user turns to your agent. Defaults to
gpt-4o-mini-tts. - Speech-to-Text (STT): the model used to transcribe what your agent says.
By default, Confident AI simulates each user turn as a cascade: transcribe your agent with the STT model, generate the next user turn with the simulation model, then speak it with the TTS model.
Speech-to-Speech
Turn on Use real-time architecture under Speech-to-Speech (S2S) to replace that cascade with a single real-time caller. This gives you lower latency and more natural interruptions, closer to how a real caller behaves.
The real-time caller runs on fal. Pick Confident AI as the provider to use Confident AI's key, or pick fal and add your own fal API key under Model Settings → Credentials.
Running a Voice Simulation
There's nothing extra to set up at run time. Run a multi-turn evaluation on your dataset and select your voice connection. Each golden becomes one call.
While a test case is generating, its status shows where the call is:
- Preparing caller...: the real-time caller is starting up (Speech-to-Speech only)
- Dialing...: Confident AI is placing the call
- In call...: the conversation is in progress
Your agent always speaks first. If a conversational golden's turns start with an assistant turn, Confident AI skips it and waits for your agent's greeting instead.
Once a test case finishes, open it to see:
- Call Recording: the full recording of the call
- Turn audio: each turn has its own audio clip alongside its transcript
- Answered in: how long your agent took to respond to each turn
- Interrupted: shown on turns where one side talked over the other
Next Steps
With your voice connection set up, run your first simulation.
Multi-Turn Evals
Run multi-turn evaluations against your voice connection.
LiveKit Integration
Trace your LiveKit voice agent in production.
Last updated on