Launch Week 3: Five days of launches
Blog

Introducing Jev-as-a-Judge: Faster, Cheaper, More Consistent Evals on Confident AI

Introducing Jev-as-a-Judge: Faster, Cheaper, More Consistent Evals on Confident AI

Most teams evaluate a tiny fraction of their production traffic.

Not because they want to. Every online eval on a trace and every sentiment label on a conversation is a judge call, and judge calls cost money. So sample rates get set to 1%, and the other 99% of production goes unscored and unclassified.

That creates a ceiling on everything built on top. A regression that affects one segment of traffic may never show up in a 1% sample. A correlation between an issue and negative sentiment needs enough labeled traces to be visible. A customer with a bad week may have only a handful of classified conversations.

So today, for Day 5 of Launch Week 03, we're launching Jev-as-a-Judge on Confident AI. It makes judge decisions fast and cheap enough that you can stop rationing them.

Stop sampling evals on 1% of live traffic

With Jev-as-a-Judge, the workflows you already run on Confident AI can cover far more of your live traffic:

  • More online evals on traces: score 20%, 50%, or more of production instead of a thin slice
  • More sentiment analysis on traces and threads: classify most conversations, not a handful
  • More classification across every dimension: issues, use cases, and safety labels on the traffic that matters

Same evaluation rules, same classifiers, same workflows. You just turn the sample rate up.

More coverage, more insights

Every insight on Confident AI is limited by how much of your traffic has been scored and labeled. Raise the coverage and every downstream workflow gets sharper:

  • Drift has more evaluated traces per agent configuration, so regressions reach significance sooner and smaller segments of live traffic become visible
  • Category Correlation has more labeled traces and threads, so relationships between issues, sentiment, and use cases stand out from the noise
  • Customer Monitoring has a complete picture of each account's sentiment and issues, not a sample of it
  • Signals catches spikes in classifier labels earlier, because more of the traffic behind them is classified

The insights no longer depend on whether the right trace happened to land in the sample.

What is Jev?

Jev is TypeSafe AI's System One model. It is not a language model. You give it some state and a bounded question, such as whether a response is supported or which label fits, and it returns a decision with a probability.

Most of an eval is a decision like that, not a writing task. Jev takes over those decisions while the LLM keeps the language work, which makes each judgment:

  • Cheap: $0.042 per million input tokens, with no charge for output tokens
  • Fast: most decisions complete in roughly 100 ms
  • Consistent: the same trace gets the same verdict, so trends reflect your agent, not judge noise

The tradeoff: because Jev is not a chat model, it does not write out its reasoning. You get the verdict and its probability, not an explanation.

Get started

Jev-as-a-Judge is live on Confident AI now. Open Project Settings → Model Settings → Decision Models, select TypeSafe as the provider and jev-latest as the model, then raise the sample rates you have been keeping low.

Jev is also available in DeepEval, our open-source evaluation framework. For the technical walkthrough, read Introducing Jev in DeepEval.

Book a demo with the Confident AI team to see Jev-as-a-Judge in action.


Do you want to brainstorm how to evaluate your LLM (application)? Ask us anything in our discord. I might give you an "aha!" moment, who knows?

Standardize AI Quality for the entire org, not just individual teams

Give all AI use cases the same quality bar with all-in-one evals, observability, and red teaming, and enforce them at scale.

AI evals for product teams, not just engineers.
Observability for production traffic.
Red teaming for security and safety.
AI governance for multiple projects at once.

More stories from us...