Launch Week 3: Five days of launches

LLM Evaluation Metrics

Overview of metrics in Confident AI

Overview

Metrics are the foundation of LLM evaluation on Confident AI. They define the criteria used to score and assess your LLM outputs — whether you're testing in development, running experiments, or monitoring production systems.

Confident AI provides two categories of metrics:

  • Pre-built metrics — Battle-tested metrics for common evaluation scenarios like answer relevancy, faithfulness, hallucination detection, and more
  • Custom metrics — Create your own metrics tailored to your specific use case using G-Eval (natural language criteria) or Code-Evals (Python code)

Both categories support single-turn (individual LLM interactions) and multi-turn (conversational) evaluations.

How Metrics Work

Metrics on Confident AI follow a simple pattern:

  1. Define your metrics — Choose from pre-built metrics or create custom ones
  2. Group into collections — Add metrics to a metric collection with specific settings (threshold, strictness, etc.)
  3. Run evaluations — Use the collection for test runs, experiments, or production monitoring
  4. Analyze results — View scores, reasoning, and pass/fail status in the dashboard

Pre-built Metrics

Confident AI offers a comprehensive library of pre-built metrics powered by LLM-as-a-judge:

MetricDescription
Answer RelevancyMeasures how relevant the response is to the input query
Argument CorrectnessJudges the arguments generated for each tool call
BiasDetects biased content in responses
Contextual PrecisionEvaluates retrieval ranking quality
Contextual RecallMeasures retrieval completeness
Contextual RelevancyAssesses relevance of retrieved context
Exact MatchChecks the output against the expected output exactly
FaithfulnessChecks if the response is grounded in the provided context
HallucinationDetects fabricated or unsupported information
Image CoherenceMeasures how well images complement the text around them
Image EditingEvaluates whether an image was edited as instructed
Image HelpfulnessMeasures how much images aid comprehension of the text
Image ReferenceEvaluates how accurately text refers to the images
MisuseDetects out-of-scope requests answered by a domain chatbot
Non-AdviceDetects unlicensed professional advice
Pattern MatchChecks the output against a regular expression
PII LeakageDetects exposure of personally identifiable information
Plan AdherenceChecks if the agent followed the plan it laid out
Plan QualityEvaluates whether the agent's plan suits the task
Prompt AlignmentChecks if the output follows your prompt instructions
Role ViolationDetects breaks in the assigned character
Step EfficiencyDetects steps the agent did not need to take
SummarizationEvaluates summary quality and accuracy
Task CompletionChecks if the task was successfully completed
Tool CorrectnessValidates correct tool/function usage
ToxicityIdentifies toxic or harmful content

Custom Metrics

When pre-built metrics don't fit your use case, create custom metrics:

Next Steps

Ready to start evaluating? Here's where to go next:

Scaling beyond prototype?For teams evaluating Confident AI in productionTalk to us

Last updated on

Built byConfident AI