Launch Week 3: Five days of launches

Audio in Traces

Add audio to your messages and play it back on each message in a trace

Overview

Confident AI lets you attach audio to your messages array in your traces and play it back right from the trace. This is especially useful for voice agents, where the transcript alone doesn't tell you how something was said, such as whether the user was cut off, the speech-to-text misheard them, or the agent's reply sounded wrong.

Add audio using Media on messages at the span or trace level. Unlike images and PDFs, audio isn't placed inside a string. It goes on a message, next to its role and content, and the trace shows a player on that message.

Add Audio to Messages

To add audio, pass a Media under any key of a message, such as audio, voice, or recording. Confident AI detects the audio automatically by its media type, and shows a player on that message:

main.py
from confident_trace import Media, init, span, update_span, shutdown

init()

@span(type="agent")
def voice_turn(question: str, answer: str):
    ...
    update_span(
        input={
            "role": "user",
            "content": "What's my balance?",
            "audio": Media.from_file(question),
        },
        output={
            "role": "assistant",
            "content": "Your balance is $42.",
            "audio": Media.from_file(answer),
        },
    )

try:
    voice_turn("question.wav", "answer.mp3")
finally:
    shutdown()

Here, both messages keep their text in content, so evaluations read them as usual, and the span shows a player on each one.

confident-trace reads the media type from the file extension for .mp3, .wav, .m4a, .ogg, .opus, .flac, and .aac files. For any other format, pass the media type yourself, such as Media.from_file(path, "audio/webm").

How Audio Is Stored

When a trace arrives, Confident AI uploads each audio file and keeps a reference to it on the message, under the same key you used:

{
  "role": "user",
  "content": "What's my balance?",
  "audio": { "mimeType": "audio/wav", "url": "https://…/audio/7c1e….wav" }
}

Only mimeType and url are kept on the audio object, so any other field you put inside it, such as a transcript, is dropped. Keep the text of the message in content instead. Audio you pass with an external URL isn't uploaded, and keeps that URL.

Add Audio to a Conversation

For multi-turn apps, you'll often want the whole conversation on the trace, with a player on every turn. Set a list of messages as the trace input, with audio attached using Media, and each message that has audio gets its own player:

main.py
from confident_trace import Media, init, span, update_trace, shutdown

init()

@span(type="agent")
def support_call():
    ...
    update_trace(
        input=[
            {
                "role": "user",
                "content": "What's my balance?",
                "audio": Media.from_file("turn-1-user.wav"),
            },
            {
                "role": "assistant",
                "content": "Your balance is $42.",
                "audio": Media.from_file("turn-1-agent.mp3"),
            },
            {
                "role": "user",
                "content": "When is my next payment due?",
                "audio": Media.from_file("turn-2-user.wav"),
            },
        ],
        output={
            "role": "assistant",
            "content": "Your next payment is due on the 1st.",
            "audio": Media.from_file("turn-2-agent.mp3"),
        },
    )

try:
    support_call()
finally:
    shutdown()

Pressing play on a message plays its audio and then continues through the audio of each message after it, in order, until you pause it or press the space bar. This lets you listen to a whole conversation in one go. Only one player plays at a time, and if the messages are on an LLM span, the same players also appear in the trace's Messages View.

Send Audio Without the SDK

If you send traces to Confident AI without confident-trace, use the same shape on each message, with the audio either inline as base64 under dataBase64 or as a link under url:

[
  {
    "role": "user",
    "content": "What's my balance?",
    "audio": { "mimeType": "audio/wav", "dataBase64": "UklGRiQAAABXQVZFZm10..." }
  },
  {
    "role": "assistant",
    "content": "Your balance is $42.",
    "audio": { "mimeType": "audio/mpeg", "url": "https://example.com/answer.mp3" }
  }
]

Confident AI recognizes audio by its audio/* media type under any key, in the input and output of both traces and spans. Audio sent as dataBase64 is uploaded and replaced with a url, and if an audio object has both, dataBase64 is used.

Audio in Model Calls

Integrations don't keep the inline audio in model calls, such as an OpenAI input_audio part. The span records its media type, with the content marked as omitted. Audio you pass to a model by URL is kept as a link to that URL.

To keep the audio of a model call on a trace, add it to a message yourself, as shown in Add Audio to Messages.

Next Steps

With audio on your messages, trace your voice agents end to end or connect their turns into conversations.

Ready to monitor AI in production?Connect traces, alerts, dashboards, and evals in one production workflowBook a demo

Last updated on

Built byConfident AI