OpenAI
Use Confident AI for LLM observability and evals for OpenAI
Overview
Confident AI lets you trace and evaluate OpenAI calls, whether standalone or used as a component within a larger application.
With confident-trace, Confident AI's OpenTelemetry-native tracing SDK, you keep the official OpenAI client exactly as it is — call init() once and every Chat Completions or Responses call shows up in the Observatory as an LLM span with its messages, token usage and cost, latency, and errors.
| Runtime | Requirements | Supported calls |
|---|---|---|
| Python | Python 3.10+ | Sync and async Chat Completions and Responses create, including streaming |
| TypeScript | Node.js 22+, openai >=7.10.0 <8 | Chat Completions and Responses create, including streaming |
Auto-Instrument
Install Dependencies
Run the following command to install
confident-tracealongside the OpenAI SDK:pip install confident-trace openaitsxis only needed if you run TypeScript source directly.npm install confident-trace 'openai@>=7.10.0 <8' npm install -D tsxyarn add confident-trace 'openai@>=7.10.0 <8' yarn add -D tsxSet Your API Keys
Get your Confident AI Project API key and set it as an environment variable, along with your OpenAI key:
export CONFIDENT_API_KEY="<your-confident-project-key>" export OPENAI_API_KEY="<your-openai-key>"Instrument OpenAI
Call
init()once before making model calls. It detects the installed OpenAI SDK and instruments it for you — keep importing your client fromopenaias usual, no wrapper needed.main.py from confident_trace import init, shutdown from openai import OpenAI init() client = OpenAI() try: response = client.responses.create( model="gpt-4.1-mini", input="Explain OpenTelemetry in one sentence.", ) print(response.output_text) finally: shutdown()src/index.ts import OpenAI from "openai"; import { init } from "confident-trace"; const runtime = init(); const client = new OpenAI(); try { const response = await client.responses.create({ model: "gpt-4.1-mini", input: "Explain OpenTelemetry in one sentence.", }); console.log(response.output_text); } finally { await runtime.shutdown(); }TypeScript needs one more thing: launch your entry point with the
confident-trace/registerpreload so the SDK can hook theopenaipackage as Node loads it.init()handles export, the preload handles instrumentation — you need both.Run OpenAI
Run your script to send the trace to Confident AI:
python main.py# Running TypeScript source directly node --import tsx --import confident-trace/register src/index.ts # Running compiled JavaScript node --import confident-trace/register dist/index.jsTo make this your normal startup command, add it to your
package.jsonscripts:package.json { "scripts": { "start": "node --import confident-trace/register dist/index.js", "dev": "node --import tsx --import confident-trace/register src/index.ts" } }Done ✅. Open the Observatory in your Confident AI project and you'll find a trace with an LLM span inside it.
What Gets Captured
Each supported model call becomes an LLM span. If there's already an active span (for example one you created with span), the call nests under it; a call with no parent starts a new trace of its own.
- Model and response details — requested model, response ID, timing, status, and token usage.
- Messages — input/output messages, finish reasons, and tool-call data returned by the model.
- Streaming output — recorded as your app consumes the stream, without reading ahead of it.
Captured content follows the content policy. Size limits are disabled by default, but you can configure a limit or redact content before export.
Chat Completions, Streaming, and Async
The quickstart used the Responses API, but every supported call is traced the same way. These examples continue after init() and client setup from the quickstart, and before shutdown().
completion = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the weather in France?"},
],
)
print(completion.choices[0].message.content)const completion = await client.chat.completions.create({
model: "gpt-4.1-mini",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "What is the weather in France?" },
],
});
console.log(completion.choices[0].message.content);with client.responses.create(
model="gpt-4.1-mini", input="Write a short poem.", stream=True
) as stream:
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)const stream = await client.responses.create({
model: "gpt-4.1-mini", input: "Write a short poem.", stream: true,
});
try {
for await (const event of stream) {
if (event.type === "response.output_text.delta") {
process.stdout.write(event.delta);
}
}
} finally {
stream.controller.abort();
}Use the official AsyncOpenAI client with the same init() setup. Await the call as usual, and consume streams with async for:
import asyncio
from openai import AsyncOpenAI
async_client = AsyncOpenAI()
async def generate_response(input: str) -> str:
response = await async_client.responses.create(
model="gpt-4.1-mini",
instructions="You are a helpful assistant.",
input=input,
)
return response.output_text
print(asyncio.run(generate_response("What is the weather in France?")))Set Trace Span Properties
Use a trace context to add properties you know before the call starts. It creates no extra span; the trace started by client.responses.create() inherits the tags, metadata, user ID, and customer ID.
from confident_trace import init, trace_context
from openai import OpenAI
init()
client = OpenAI()
with trace_context(
tags=["support"],
metadata={"release": "2026-09"},
user_id="user-42",
customer_id="customer-7",
):
response = client.responses.create(
model="gpt-4.1-mini",
input="Explain OpenTelemetry in one sentence.",
)import { init, traceContext } from "confident-trace";
import OpenAI from "openai";
init();
const client = new OpenAI();
const response = await traceContext(
{
tags: ["support"],
metadata: { release: "2026-09" },
userId: "user-42",
customerId: "customer-7",
},
() => client.responses.create({
model: "gpt-4.1-mini",
input: "Explain OpenTelemetry in one sentence.",
}),
);See users and customers to set an optional display name alongside each ID, and trace context for every supported trace property and update behavior.
Instrumenting Multi-Turn
You do not need turn() when one OpenAI entry-point call is already one conversational turn—the integration creates that turn's trace automatically. Use turn() when you want to define the boundary yourself, such as grouping two sequential OpenAI calls into one turn. Reuse the same thread ID on later turns to group them into one conversation.
from confident_trace import init, turn
init()
with turn("support-turn", thread_id="chat-42"):
context = client.responses.create(model="gpt-4.1-mini", input="Find the relevant account details.")
answer = client.responses.create(model="gpt-4.1-mini", input=f"Summarize these details: {context.output_text}")import { init, turn } from "confident-trace";
init();
const answer = await turn({ name: "support-turn", threadId: "chat-42" }, async () => {
const context = await client.responses.create({ model: "gpt-4.1-mini", input: "Find the relevant account details." });
return client.responses.create({ model: "gpt-4.1-mini", input: `Summarize these details: ${context.output_text}` });
});See threads for thread I/O, turn IDs, and user IDs.
Troubleshooting
- No spans: make sure
init()runs before your first model call, and that the process reachesshutdown()so buffered spans are flushed. - Missing final stream content: fully consume or close the stream before
shutdown(). - Duplicate spans: you have two instrumentors on the same client. Don't wrap the client with another provider instrumentor while
confident-traceis active; passinit(instrumentations=())if external instrumentation already supplies your OpenAI spans. - Missing content: check masking and content controls, truncation limits, and whether the API you called is in the supported scope above. Binary multimodal payloads are omitted.
- No spans: make sure
init()runs before your first model call, that your start command includes--import confident-trace/register, and that the process reachesshutdown()so buffered spans are flushed.runtime.getInstrumentationStatus()tells you whether the OpenAI hook attached. - Missing final stream content: fully consume or abort the stream before
shutdown(). - Duplicate spans: you have two instrumentors on the same client. Don't wrap the client with another provider instrumentor while
confident-traceis active; passinit({ instrumentations: [] })if external instrumentation already supplies your OpenAI spans. - Missing content: check masking and content controls, truncation limits, and whether the API you called is in the supported scope above. Binary multimodal payloads are omitted.
For general setup issues, see troubleshooting.
Disable OpenAI Instrumentation
Pass init() a list of integration identifiers to opt in to only those integrations. The identifier for OpenAI is "openai" in Python and TypeScript; omit it to disable this integration. An empty list disables all automatic instrumentation:
from confident_trace import init
init(instrumentations=())
# Use ("openai",) to opt in; omit "openai" to disable it.import { init } from "confident-trace";
init({ instrumentations: [] });
// Use ["openai"] to opt in; omit "openai" to disable it.This turns off Confident AI's automatic instrumentation; calls made after initialization are not instrumented by this integration.
Next Steps
Online Evals
Run evaluations on traces and spans in real-time as they're ingested into Confident AI to monitor AI quality in production.
Threads
Group multi-turn conversations into threads, set turn I/O, and evaluate entire conversations as a single unit.
Last updated on