LLM Evaluation Metrics
Overview of metrics in Confident AI
Overview
Metrics are the foundation of LLM evaluation on Confident AI. They define the criteria used to score and assess your LLM outputs — whether you're testing in development, running experiments, or monitoring production systems.
Confident AI provides two categories of metrics:
- Pre-built metrics — Battle-tested metrics for common evaluation scenarios like answer relevancy, faithfulness, hallucination detection, and more
- Custom metrics — Create your own metrics tailored to your specific use case using G-Eval (natural language criteria) or Code-Evals (Python code)
Both categories support single-turn (individual LLM interactions) and multi-turn (conversational) evaluations.
How Metrics Work
Metrics on Confident AI follow a simple pattern:
- Define your metrics — Choose from pre-built metrics or create custom ones
- Group into collections — Add metrics to a metric collection with specific settings (threshold, strictness, etc.)
- Run evaluations — Use the collection for test runs, experiments, or production monitoring
- Analyze results — View scores, reasoning, and pass/fail status in the dashboard
Pre-built Metrics
Confident AI offers a comprehensive library of pre-built metrics powered by LLM-as-a-judge:
| Metric | Description |
|---|---|
| Answer Relevancy | Measures how relevant the response is to the input query |
| Argument Correctness | Judges the arguments generated for each tool call |
| Bias | Detects biased content in responses |
| Contextual Precision | Evaluates retrieval ranking quality |
| Contextual Recall | Measures retrieval completeness |
| Contextual Relevancy | Assesses relevance of retrieved context |
| Exact Match | Checks the output against the expected output exactly |
| Faithfulness | Checks if the response is grounded in the provided context |
| Hallucination | Detects fabricated or unsupported information |
| Image Coherence | Measures how well images complement the text around them |
| Image Editing | Evaluates whether an image was edited as instructed |
| Image Helpfulness | Measures how much images aid comprehension of the text |
| Image Reference | Evaluates how accurately text refers to the images |
| Misuse | Detects out-of-scope requests answered by a domain chatbot |
| Non-Advice | Detects unlicensed professional advice |
| Pattern Match | Checks the output against a regular expression |
| PII Leakage | Detects exposure of personally identifiable information |
| Plan Adherence | Checks if the agent followed the plan it laid out |
| Plan Quality | Evaluates whether the agent's plan suits the task |
| Prompt Alignment | Checks if the output follows your prompt instructions |
| Role Violation | Detects breaks in the assigned character |
| Step Efficiency | Detects steps the agent did not need to take |
| Summarization | Evaluates summary quality and accuracy |
| Task Completion | Checks if the task was successfully completed |
| Tool Correctness | Validates correct tool/function usage |
| Toxicity | Identifies toxic or harmful content |
| Metric | Description |
|---|---|
| Conversation Completeness | Measures if the conversation achieved its goal |
| Goal Accuracy | Measures if the agent reached the user's goal |
| Knowledge Retention | Checks if context is maintained across turns |
| Role Adherence | Evaluates consistency with assigned persona |
| Topic Adherence | Checks if the chatbot stays on its relevant topics |
| Turn Contextual Precision | Evaluates retrieval ranking quality at each turn |
| Turn Contextual Recall | Measures retrieval completeness at each turn |
| Turn Contextual Relevancy | Assesses relevance of retrieved context at each turn |
| Turn Faithfulness | Checks if each turn is grounded in its retrieved context |
| Turn Relevancy | Assesses relevance of each conversational turn |
Custom Metrics
When pre-built metrics don't fit your use case, create custom metrics:
G-Eval
Define evaluation criteria in natural language. Best for subjective qualities like tone, helpfulness, or domain-specific correctness.
Code-Evals
Write Python code directly on Confident AI. Best for deterministic checks, format validation, or complex calculations.
Arena G-Eval
Pick the better of two or more versions of your LLM app. Best for comparing prompts or models against each other.
DAG
Build a deterministic decision tree for evaluation. Best for conditional rules, gates, and known scoring paths.
Prompt
Write your own judge prompt and fill it with data from your traces. Best when you already have a judge prompt that works.
Next Steps
Ready to start evaluating? Here's where to go next:
Metric Collections
Learn how to group metrics and configure settings for remote evaluations.
Custom Metrics
Create metrics tailored to your specific evaluation needs.
Last updated on