Suppose you want to answer one question: Which agent issues actually cause negative user sentiment?
Today, you would start in your traces. Filter for one issue AND negative sentiment. Count the results. Clear the filters, choose another issue, and repeat — across every possible combination of issue, sentiment, use case, and safety signal.
Even after all that work, a raw count does not tell you whether two labels appear together unusually often or whether you are simply looking at a common issue.
In other words, answering a straightforward product question requires a long, manual analysis.
So today, for Day 3 of Launch Week 03, we're launching Category Correlation on Confident AI. Run classifiers on incoming traces, and the Correlations page automatically shows which categories move together, which combinations appear more or less often than expected, and the traces behind every pattern.
Stop testing every combination by hand
Filters are useful when you already know what you are looking for. They are a poor discovery tool.
If you have ten issue labels and three sentiment labels, that is already 30 combinations to test. Add use cases, security vulnerabilities, and other classifiers, and the search space grows faster than any team can reasonably investigate.
Worse, manual filtering only answers the combination you thought to ask about. The strongest relationship in your production traffic might be one you never put into the filter bar.
Category Correlation analyzes every combination for you. Instead of constructing dozens of AND statements, choose two classifiers — such as Issues and Sentiment — and see the complete relationship between their labels in one view.
Classify live traces once
The Correlations page is powered by the classifiers already running on Confident AI.
Create classifiers for the dimensions that matter to your product:
- Issues — resolution failures, incorrect answers, broken tool calls, or any failure taxonomy your team uses
- Sentiment — positive, negative, neutral, or a more specific customer reaction
- Use cases — support resolution, analytics, booking, retrieval, or other jobs your agent performs
- Security and safety — policy violations, risky behavior, or vulnerability categories
As incoming traces are classified, their labels become signals that can be analyzed together. You define the categories once; Confident AI continuously builds the relationship map from live production traffic.
See which categories actually move together
Select one classifier for each axis and the Correlations page does the analysis that filters cannot.
An overall Cramér's V score measures the strength of the relationship between the two classifiers. The heatmap then breaks that relationship down label by label:
- Appear together more — the combination occurs more often than expected
- Appear together less — the combination occurs less often than expected
- No meaningful link — there is not enough evidence of a relationship
This matters because frequency alone can be misleading. A common issue may appear alongside negative sentiment many times simply because it appears everywhere. Category Correlation highlights the pairings that are unusually concentrated, not just the labels with the largest totals.
Correlation does not, by itself, prove causation. It does tell you where the evidence is strongest — turning "what might be driving negative sentiment?" from a filtering exercise into a focused investigation.
From a pattern to the exact traces
A correlation is only useful if you can inspect what produced it.
Click any cell in the heatmap to open the traces behind that pairing. If Support Resolution appears with Negative Sentiment far more often than expected, you can move directly from the pattern to the conversations, tool calls, and outputs that explain it.
That gives every team a faster path from production signal to action:
- Product teams see which issues have the strongest relationship with user outcomes
- Engineers get the exact traces needed to diagnose the behavior
- Evaluation teams learn which failure modes deserve new test cases and metrics
You no longer have to search through every combination hoping to find the important one. The important relationships surface themselves.
This also works for threads (conversations)
Not every useful signal belongs to one trace. Sentiment, resolution, escalation, repetition, and other outcomes often only make sense across a full multi-turn conversation.
Category Correlation also works with threads, so you can run classifiers over complete conversations and discover which thread-level categories appear together. The workflow stays the same: classify the conversations, choose two dimensions, and inspect the relationships and underlying threads that matter.
Get started
Category Correlation is live on Confident AI now.
Create classifiers for the production signals you care about, let them run on incoming traces, and discover which categories are connected. Book a demo with the Confident AI team to see Category Correlation in action.
Do you want to brainstorm how to evaluate your LLM (application)? Ask us anything in our discord. I might give you an "aha!" moment, who knows?
Standardize AI Quality for the entire org, not just individual teams
Give all AI use cases the same quality bar with all-in-one evals, observability, and red teaming, and enforce them at scale.

