Multi-Modal Tracing
Capture images, PDFs, and audio in your traces, from model calls or attached yourself
Overview
Multimodal AI apps don't just work with text. Document assistants read PDFs, vision agents look at screenshots and charts, and voice agents listen and talk back. Confident AI captures the images, PDFs, and audio in your traces alongside the text of your messages, so you can see exactly what your model was given:
- Debugging: See the actual image or document a model answered about, rather than a filename or a truncated base64 string
- Voice Agents: Listen to what was said on each message of a conversation, right next to its transcript
Media gets onto a trace in two ways. Supported integrations capture the images and PDFs your app sends to a model automatically, and you can also add media to any trace or span manually using Media.
Supported Media
Confident AI supports three types of media, and displays each one in the way that suits it best:
Images
Capture the images sent to your models, and add images to any trace or span.
PDFs
Capture the PDFs sent to your models, and add documents to any trace or span.
Audio
Add audio to your messages and play it back on each message in a trace.
How It Works
Integrations record each media part in a model call's messages in one of three ways, depending on how you passed it to the model and what type of file it is:
| You send the model | Recorded on the span as |
|---|---|
| An image or PDF as inline data (base64) | The file itself, viewable in the trace |
| A URL, or a cloud storage location (S3, GCS) | A reference to that location, the file is never downloaded |
| Inline audio, video, or any other file type | The media type only, with the content marked as omitted |
| A provider file ID (an uploaded file) | The media type only, with the content marked as omitted |
Images (any image/* type) and PDFs are the only inline types kept from model calls. A URL reference is kept for any type, so an audio file you pass to the model by URL still shows up as a link to that audio file. If you want the audio itself on the trace, add it to a message manually.
Supported Integrations
The following integrations capture images and PDFs from the messages they record:
| Integration | Python | TypeScript |
|---|---|---|
| OpenAI | Yes | Yes |
| Anthropic | Yes | Yes |
| Google GenAI | Yes | Yes |
| AWS Bedrock | Yes | — |
| LiteLLM | Yes | — |
| OpenRouter | Yes | Yes |
| Portkey | Yes | Yes |
| LangChain and LangGraph | Yes | Yes |
| Vercel AI SDK | — | Yes |
| Mastra | — | Yes |
Frameworks that don't record model calls themselves, such as LlamaIndex, CrewAI, Agno, and smolagents, capture media whenever the provider they call is one of the integrations above.
Create Media
For media that isn't part of a model call, such as a file your app loads, a chart a tool generated, or the recording of a voice turn, wrap it in Media and pass it to update_span() / updateSpan() or update_trace() / updateTrace(). You can create a Media from wherever the file is:
| Source | Python | TypeScript |
|---|---|---|
| A local file | Media.from_file("chart.png") | Media.fromFile("chart.png") |
| A URL | Media(uri="https://…/chart.png") | new Media({ uri: "https://…/chart.png" }) |
| Bytes | Media.from_bytes(data, "image/png") | Media.fromBytes(data, "image/png") |
| Base64 | Media.from_base64(text, "image/png") | Media.fromBase64(text, "image/png") |
If you don't pass a media type, confident-trace reads it from the file extension. A URL is recorded as a reference and never downloaded, while a local file, bytes, or base64 is exported with the span. Media that can't be kept, such as a file over the size limits, a file that can't be read, or a type other than an image, PDF, or audio, is replaced with a note saying it wasn't captured.
Where a Media can go depends on its type. Images and PDFs can be a value on their own or part of a string, while audio goes on a message. Each type's page walks through it with an example:
Media Size Limits
To keep exports bounded, confident-trace limits how much inline media each span carries, whether captured from a model call or added yourself. References (URLs and storage locations) are never counted toward these limits.
| Limit | Python init() | TypeScript init() | Default | What it does |
|---|---|---|---|---|
| Per media item | max_media_bytes | maxMediaBytes | 5 MiB | The largest single image, PDF, or audio file that is kept, measured in decoded bytes |
| Per span | — | maxMediaTotalBytes | 16 MiB | The total inline media kept across all of a span's messages and added media |
A media item over either limit is omitted, never truncated. In a model call's messages, it stays as a media part with its type and an omission marker, so the span still shows that the model was sent a file. Media you added yourself is replaced with a note instead. Within a span, media items are kept in the order they appear until the per-span budget runs out.
from confident_trace import init
init(max_media_bytes=10 * 1024 * 1024)import { init } from "confident-trace";
const runtime = init({
maxMediaBytes: 10 * 1024 * 1024,
maxMediaTotalBytes: 32 * 1024 * 1024,
});Disable Media Capture
To keep text content but stop exporting inline media, set max_media_bytes / maxMediaBytes to 0. Every inline image, PDF, and audio file is then left out, while URL references are still recorded.
from confident_trace import init
init(max_media_bytes=0)import { init } from "confident-trace";
const runtime = init({ maxMediaBytes: 0 });Next Steps
Now that you know how media is captured, see how each type works in detail.
Images
Capture the images sent to your models, and add images to any trace or span.
PDFs
Capture the PDFs sent to your models, and add documents to any trace or span.
Audio
Add audio to your messages and play it back on each message in a trace.
Last updated on