Launch Week 3: Five days of launches

OpenAI

Use Confident AI for LLM observability and evals for OpenAI

Overview

Confident AI lets you trace and evaluate OpenAI calls, whether standalone or used as a component within a larger application.

With confident-trace, Confident AI's OpenTelemetry-native tracing SDK, you keep the official OpenAI client exactly as it is — call init() once and every Chat Completions or Responses call shows up in the Observatory as an LLM span with its messages, token usage and cost, latency, and errors.

RuntimeRequirementsSupported calls
PythonPython 3.10+Sync and async Chat Completions and Responses create, including streaming
TypeScriptNode.js 22+, openai >=7.10.0 <8Chat Completions and Responses create, including streaming

Auto-Instrument

  1. Install Dependencies

    Run the following command to install confident-trace alongside the OpenAI SDK:

    pip install confident-trace openai
  2. Set Your API Keys

    Get your Confident AI Project API key and set it as an environment variable, along with your OpenAI key:

    export CONFIDENT_API_KEY="<your-confident-project-key>"
    export OPENAI_API_KEY="<your-openai-key>"
  3. Instrument OpenAI

    Call init() once before making model calls. It detects the installed OpenAI SDK and instruments it for you — keep importing your client from openai as usual, no wrapper needed.

    main.py
    from confident_trace import init, shutdown
    from openai import OpenAI
    
    init()
    client = OpenAI()
    
    try:
        response = client.responses.create(
            model="gpt-4.1-mini",
            input="Explain OpenTelemetry in one sentence.",
        )
        print(response.output_text)
    finally:
        shutdown()
  4. Run OpenAI

    Run your script to send the trace to Confident AI:

    python main.py

    Done ✅. Open the Observatory in your Confident AI project and you'll find a trace with an LLM span inside it.

What Gets Captured

Each supported model call becomes an LLM span. If there's already an active span (for example one you created with span), the call nests under it; a call with no parent starts a new trace of its own.

  • Model and response details — requested model, response ID, timing, status, and token usage.
  • Messages — input/output messages, finish reasons, and tool-call data returned by the model.
  • Streaming output — recorded as your app consumes the stream, without reading ahead of it.

Captured content follows the content policy. Size limits are disabled by default, but you can configure a limit or redact content before export.

Chat Completions, Streaming, and Async

The quickstart used the Responses API, but every supported call is traced the same way. These examples continue after init() and client setup from the quickstart, and before shutdown().

completion = client.chat.completions.create(
    model="gpt-4.1-mini",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is the weather in France?"},
    ],
)
print(completion.choices[0].message.content)

Set Trace Span Properties

Use a trace context to add properties you know before the call starts. It creates no extra span; the trace started by client.responses.create() inherits the tags, metadata, user ID, and customer ID.

main.py
from confident_trace import init, trace_context
from openai import OpenAI

init()

client = OpenAI()

with trace_context(
    tags=["support"],
    metadata={"release": "2026-09"},
    user_id="user-42",
    customer_id="customer-7",
):
    response = client.responses.create(
        model="gpt-4.1-mini",
        input="Explain OpenTelemetry in one sentence.",
    )

See users and customers to set an optional display name alongside each ID, and trace context for every supported trace property and update behavior.

Instrumenting Multi-Turn

You do not need turn() when one OpenAI entry-point call is already one conversational turn—the integration creates that turn's trace automatically. Use turn() when you want to define the boundary yourself, such as grouping two sequential OpenAI calls into one turn. Reuse the same thread ID on later turns to group them into one conversation.

main.py
from confident_trace import init, turn

init()

with turn("support-turn", thread_id="chat-42"):
    context = client.responses.create(model="gpt-4.1-mini", input="Find the relevant account details.")
    answer = client.responses.create(model="gpt-4.1-mini", input=f"Summarize these details: {context.output_text}")

See threads for thread I/O, turn IDs, and user IDs.

Troubleshooting

  • No spans: make sure init() runs before your first model call, and that the process reaches shutdown() so buffered spans are flushed.
  • Missing final stream content: fully consume or close the stream before shutdown().
  • Duplicate spans: you have two instrumentors on the same client. Don't wrap the client with another provider instrumentor while confident-trace is active; pass init(instrumentations=()) if external instrumentation already supplies your OpenAI spans.
  • Missing content: check masking and content controls, truncation limits, and whether the API you called is in the supported scope above. Binary multimodal payloads are omitted.

For general setup issues, see troubleshooting.

Disable OpenAI Instrumentation

Pass init() a list of integration identifiers to opt in to only those integrations. The identifier for OpenAI is "openai" in Python and TypeScript; omit it to disable this integration. An empty list disables all automatic instrumentation:

main.py
from confident_trace import init
init(instrumentations=())
# Use ("openai",) to opt in; omit "openai" to disable it.

This turns off Confident AI's automatic instrumentation; calls made after initialization are not instrumented by this integration.

Next Steps

Need help instrumenting your application?Connect your model calls and agent workflows to Confident AITalk to an expert

Last updated on

Built byConfident AI