Launch Week 3: Five days of launches

Multi-Modal Tracing

Capture images, PDFs, and audio in your traces, from model calls or attached yourself

Overview

Multimodal AI apps don't just work with text. Document assistants read PDFs, vision agents look at screenshots and charts, and voice agents listen and talk back. Confident AI captures the images, PDFs, and audio in your traces alongside the text of your messages, so you can see exactly what your model was given:

  • Debugging: See the actual image or document a model answered about, rather than a filename or a truncated base64 string
  • Voice Agents: Listen to what was said on each message of a conversation, right next to its transcript

Media gets onto a trace in two ways. Supported integrations capture the images and PDFs your app sends to a model automatically, and you can also add media to any trace or span manually using Media.

Supported Media

Confident AI supports three types of media, and displays each one in the way that suits it best:

How It Works

Integrations record each media part in a model call's messages in one of three ways, depending on how you passed it to the model and what type of file it is:

You send the modelRecorded on the span as
An image or PDF as inline data (base64)The file itself, viewable in the trace
A URL, or a cloud storage location (S3, GCS)A reference to that location, the file is never downloaded
Inline audio, video, or any other file typeThe media type only, with the content marked as omitted
A provider file ID (an uploaded file)The media type only, with the content marked as omitted

Images (any image/* type) and PDFs are the only inline types kept from model calls. A URL reference is kept for any type, so an audio file you pass to the model by URL still shows up as a link to that audio file. If you want the audio itself on the trace, add it to a message manually.

Supported Integrations

The following integrations capture images and PDFs from the messages they record:

IntegrationPythonTypeScript
OpenAIYesYes
AnthropicYesYes
Google GenAIYesYes
AWS BedrockYes—
LiteLLMYes—
OpenRouterYesYes
PortkeyYesYes
LangChain and LangGraphYesYes
Vercel AI SDK—Yes
Mastra—Yes

Frameworks that don't record model calls themselves, such as LlamaIndex, CrewAI, Agno, and smolagents, capture media whenever the provider they call is one of the integrations above.

Create Media

For media that isn't part of a model call, such as a file your app loads, a chart a tool generated, or the recording of a voice turn, wrap it in Media and pass it to update_span() / updateSpan() or update_trace() / updateTrace(). You can create a Media from wherever the file is:

SourcePythonTypeScript
A local fileMedia.from_file("chart.png")Media.fromFile("chart.png")
A URLMedia(uri="https://…/chart.png")new Media({ uri: "https://…/chart.png" })
BytesMedia.from_bytes(data, "image/png")Media.fromBytes(data, "image/png")
Base64Media.from_base64(text, "image/png")Media.fromBase64(text, "image/png")

If you don't pass a media type, confident-trace reads it from the file extension. A URL is recorded as a reference and never downloaded, while a local file, bytes, or base64 is exported with the span. Media that can't be kept, such as a file over the size limits, a file that can't be read, or a type other than an image, PDF, or audio, is replaced with a note saying it wasn't captured.

Where a Media can go depends on its type. Images and PDFs can be a value on their own or part of a string, while audio goes on a message. Each type's page walks through it with an example:

Media Size Limits

To keep exports bounded, confident-trace limits how much inline media each span carries, whether captured from a model call or added yourself. References (URLs and storage locations) are never counted toward these limits.

LimitPython init()TypeScript init()DefaultWhat it does
Per media itemmax_media_bytesmaxMediaBytes5 MiBThe largest single image, PDF, or audio file that is kept, measured in decoded bytes
Per span—maxMediaTotalBytes16 MiBThe total inline media kept across all of a span's messages and added media

A media item over either limit is omitted, never truncated. In a model call's messages, it stays as a media part with its type and an omission marker, so the span still shows that the model was sent a file. Media you added yourself is replaced with a note instead. Within a span, media items are kept in the order they appear until the per-span budget runs out.

main.py
from confident_trace import init

init(max_media_bytes=10 * 1024 * 1024)

Disable Media Capture

To keep text content but stop exporting inline media, set max_media_bytes / maxMediaBytes to 0. Every inline image, PDF, and audio file is then left out, while URL references are still recorded.

main.py
from confident_trace import init

init(max_media_bytes=0)

Next Steps

Now that you know how media is captured, see how each type works in detail.

Ready to monitor AI in production?Connect traces, alerts, dashboards, and evals in one production workflowBook a demo

Last updated on

Built byConfident AI