Launch Week 3: Five days of launches
Blog

Introducing AI Drift Detection: Find What Changed in Your Production Agents

Introducing AI Drift Detection: Find What Changed in Your Production Agents

Most AI observability tools give you traces, then leave you to do the investigation yourself.

An evaluator score drops. You filter by time range, compare deployments, slice by model and provider, inspect dozens of traces, and try to work out whether the problem affects everyone or one corner of your traffic. The evidence is there, but finding the insight is still a long, manual analysis.

So today, for Day 2 of Launch Week 03, we're launching the Drift page on Confident AI. It automatically finds changes in your production agents, ties them to the agent configuration that caused them, and shows where their impact is concentrated.

Instead of searching your traces for issues, the issues find you.

Start with the answer, not another dashboard

When an AI agent's behavior changes, your team needs three answers:

  1. When did it change?
  2. Which agent configuration introduced the change?
  3. Where is the change concentrated?

The Drift page answers all three automatically. It monitors evaluation scores, error rates, latency, cost, classifications, user sentiment, and more across your production traffic. When something meaningful changes, it surfaces the finding and links it back to the traces that explain it.

The investigation no longer starts with a blank filter panel. It starts with a specific finding: this agent configuration regressed, on this metric, for this part of your traffic.

Every agent configuration, identified automatically

A/B test 10 models at once. The same online evals and classifiers run across every variation, so you can compare which configuration produces the best answers — not just which one is fastest or cheapest.

Instrument your agent with confident-trace, and the Drift page keeps track of which agent configuration produced every trace. It fingerprints the OpenTelemetry-native data — including traces, spans, model calls, prompts, providers, endpoints, and tool use — then assigns each distinct configuration a version automatically. Semantic convention support keeps that fingerprint consistent across the components in each run.

Each assigned version represents one distinct agent configuration and appears on a timeline showing when it received traffic. Choose a baseline and Confident AI compares every other configuration against it across quality and operational metrics.

Because every variation is scored by the same evals and classifiers, the comparison is consistent across evaluation scores, user sentiment, failure categories, error rates, latency, and cost. The Drift page shows which configuration performs best without a separate experiment dashboard or manually reconciled exports.

Anomalies and regressions, surfaced for you

The Drift page looks for two different kinds of change:

  • Anomalies — a sudden shift in a version's behavior compared with its own recent history
  • Regressions — a sustained difference between a selected version and its baseline

This is not a simple before-and-after comparison that flags every line moving in the wrong direction. Anomaly detection uses the historical median and median absolute deviation, then requires the current bucket to cross a modified z-score threshold:

zmodified=0.6745(xx~)MAD,zmodified3.5z_{\mathrm{modified}} = \frac{0.6745\left(x - \tilde{x}\right)}{\operatorname{MAD}}, \qquad \left|z_{\mathrm{modified}}\right| \ge 3.5

Regression detection applies the same rigor across versions. Changes in mean metrics must be large enough to matter, while rate metrics must reach statistical significance at p0.05p \le 0.05. A chart moving is not enough; the evidence has to support a real change.

The result is not just a warning that a chart moved. It tells you whether the agent suddenly departed from its normal behavior or whether a new version is consistently worse than the baseline — then takes you directly to the contributing traces.

Find drift in every segment of live traffic

An overall score can stay flat while one segment of live traffic regresses.

A model update might improve common support requests while making billing conversations less accurate. It might reduce average latency while slowing down tool-heavy workflows. Aggregate performance can look healthy while a region, use case, integration, or agent configuration quietly regresses.

Whenever the Drift page finds a meaningful change, it searches for the segments where that metric moved most. Those segments can include model, provider, integration, tags, trace metadata, retrieval settings, and classifier labels.

The results are ranked by impact, balancing the size of the change with the number of traces affected. Click a segment and Confident AI opens the exact traces with the relevant filters already applied.

No segment of your live traffic gets hidden by the average.

Built for the stack you already run

The Drift page works with OpenTelemetry-native data across more than 25 integrations, including LangChain, LangGraph, Google ADK, and Amazon Bedrock AgentCore.

Your traces arrive through the same OTel pipeline regardless of framework or model provider. Confident AI versions each agent configuration, detects the changes, finds the affected segments, and gives technical and non-technical teams one place to understand what happened.

Get started

The Drift page is live on Confident AI now.

Start sending production traces and let the insights find you. Book a demo with the Confident AI team to see the Drift page in action.


Do you want to brainstorm how to evaluate your LLM (application)? Ask us anything in our discord. I might give you an "aha!" moment, who knows?

Standardize AI Quality for the entire org, not just individual teams

Give all AI use cases the same quality bar with all-in-one evals, observability, and red teaming, and enforce them at scale.

AI evals for product teams, not just engineers.
Observability for production traffic.
Red teaming for security and safety.
AI governance for multiple projects at once.

More stories from us...