Audio in Traces
Add audio to your messages and play it back on each message in a trace
Overview
Confident AI lets you attach audio to your messages array in your traces and play it back right from the trace. This is especially useful for voice agents, where the transcript alone doesn't tell you how something was said, such as whether the user was cut off, the speech-to-text misheard them, or the agent's reply sounded wrong.
Add audio using Media on messages at the span or trace level. Unlike images and PDFs, audio isn't placed inside a string. It goes on a message, next to its role and content, and the trace shows a player on that message.
Add Audio to Messages
To add audio, pass a Media under any key of a message, such as audio, voice, or recording. Confident AI detects the audio automatically by its media type, and shows a player on that message:
from confident_trace import Media, init, span, update_span, shutdown
init()
@span(type="agent")
def voice_turn(question: str, answer: str):
...
update_span(
input={
"role": "user",
"content": "What's my balance?",
"audio": Media.from_file(question),
},
output={
"role": "assistant",
"content": "Your balance is $42.",
"audio": Media.from_file(answer),
},
)
try:
voice_turn("question.wav", "answer.mp3")
finally:
shutdown()Here, both messages keep their text in content, so evaluations read them as usual, and the span shows a player on each one.
confident-trace reads the media type from the file extension for .mp3, .wav, .m4a, .ogg, .opus, .flac, and .aac files. For any other format, pass the media type yourself, such as Media.from_file(path, "audio/webm").
import { Media, init, span, updateSpan } from "confident-trace";
const runtime = init();
const voiceTurn = span({ name: "voice_turn", type: "agent" }, (question: string, answer: string) => {
...
updateSpan({
input: {
role: "user",
content: "What's my balance?",
audio: Media.fromFile(question),
},
output: {
role: "assistant",
content: "Your balance is $42.",
audio: Media.fromFile(answer),
},
});
});
try {
voiceTurn("question.wav", "answer.mp3");
} finally {
await runtime.shutdown();
}Here, both messages keep their text in content, so evaluations read them as usual, and the span shows a player on each one.
confident-trace reads the media type from the file extension for .mp3, .wav, .m4a, .ogg, .opus, .flac, and .aac files. For any other format, pass the media type yourself, such as Media.fromFile(path, "audio/webm").
How Audio Is Stored
When a trace arrives, Confident AI uploads each audio file and keeps a reference to it on the message, under the same key you used:
{
"role": "user",
"content": "What's my balance?",
"audio": { "mimeType": "audio/wav", "url": "https://…/audio/7c1e….wav" }
}Only mimeType and url are kept on the audio object, so any other field you put inside it, such as a transcript, is dropped. Keep the text of the message in content instead. Audio you pass with an external URL isn't uploaded, and keeps that URL.
Add Audio to a Conversation
For multi-turn apps, you'll often want the whole conversation on the trace, with a player on every turn. Set a list of messages as the trace input, with audio attached using Media, and each message that has audio gets its own player:
from confident_trace import Media, init, span, update_trace, shutdown
init()
@span(type="agent")
def support_call():
...
update_trace(
input=[
{
"role": "user",
"content": "What's my balance?",
"audio": Media.from_file("turn-1-user.wav"),
},
{
"role": "assistant",
"content": "Your balance is $42.",
"audio": Media.from_file("turn-1-agent.mp3"),
},
{
"role": "user",
"content": "When is my next payment due?",
"audio": Media.from_file("turn-2-user.wav"),
},
],
output={
"role": "assistant",
"content": "Your next payment is due on the 1st.",
"audio": Media.from_file("turn-2-agent.mp3"),
},
)
try:
support_call()
finally:
shutdown()import { Media, init, span, updateTrace } from "confident-trace";
const runtime = init();
const supportCall = span({ name: "support_call", type: "agent" }, () => {
...
updateTrace({
input: [
{
role: "user",
content: "What's my balance?",
audio: Media.fromFile("turn-1-user.wav"),
},
{
role: "assistant",
content: "Your balance is $42.",
audio: Media.fromFile("turn-1-agent.mp3"),
},
{
role: "user",
content: "When is my next payment due?",
audio: Media.fromFile("turn-2-user.wav"),
},
],
output: {
role: "assistant",
content: "Your next payment is due on the 1st.",
audio: Media.fromFile("turn-2-agent.mp3"),
},
});
});
try {
supportCall();
} finally {
await runtime.shutdown();
}Pressing play on a message plays its audio and then continues through the audio of each message after it, in order, until you pause it or press the space bar. This lets you listen to a whole conversation in one go. Only one player plays at a time, and if the messages are on an LLM span, the same players also appear in the trace's Messages View.
Send Audio Without the SDK
If you send traces to Confident AI without confident-trace, use the same shape on each message, with the audio either inline as base64 under dataBase64 or as a link under url:
[
{
"role": "user",
"content": "What's my balance?",
"audio": { "mimeType": "audio/wav", "dataBase64": "UklGRiQAAABXQVZFZm10..." }
},
{
"role": "assistant",
"content": "Your balance is $42.",
"audio": { "mimeType": "audio/mpeg", "url": "https://example.com/answer.mp3" }
}
]Confident AI recognizes audio by its audio/* media type under any key, in the input and output of both traces and spans. Audio sent as dataBase64 is uploaded and replaced with a url, and if an audio object has both, dataBase64 is used.
Audio in Model Calls
Integrations don't keep the inline audio in model calls, such as an OpenAI input_audio part. The span records its media type, with the content marked as omitted. Audio you pass to a model by URL is kept as a link to that URL.
To keep the audio of a model call on a trace, add it to a message yourself, as shown in Add Audio to Messages.
Next Steps
With audio on your messages, trace your voice agents end to end or connect their turns into conversations.
LiveKit
Trace LiveKit voice agents, with call transcripts and recordings on every trace.
Thread Traces
Group traces into threads to track multi-turn conversations and evaluate entire workflows.
Last updated on