Metrics
Overview
The Confident AI SDK exposes every Metric method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
List Metrics
Lists all the custom metrics in your Confident AI project.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.metrics.list()For async mode, call a_list and await it as shown below:
result = await client.metrics.a_list(...)import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.metrics.list();Returns
This method returns an object of type MetricList.
Create Metric
Creates a custom metric in your Confident AI project and returns it. A GEVAL metric scores against criteria or evaluationSteps; a DAG metric needs algorithm set to DAG and a dag.
from confident_ai import ConfidentAI
from confident_ai.common import MetricAlgorithm
from confident_ai.common import MetricDag
from confident_ai.common import Rubric
client = ConfidentAI()
result = client.metrics.create(
name="Correctness",
multi_turn=False,
criteria="Determine if the actual output is correct based on the expected output.",
evaluation_steps=[
"Compare the actual output with the expected output.",
"Penalise any factual contradiction."
],
evaluation_params=["actualOutput", "expectedOutput"],
rubric=[
Rubric(
score_range=[8, 10],
expected_outcome="The answer is factually correct and complete."
)
],
algorithm=MetricAlgorithm.DEFAULT,
dag=MetricDag(
nodes={
"root": {
"type": "BinaryJudgementNode",
"criteria": "Does the actual output answer the input?",
"evaluation_params": ["input", "actual_output"],
"children": ["pass", "fail"]
},
"pass": {"type": "VerdictNode", "verdict": True, "score": 10},
"fail": {"type": "VerdictNode", "verdict": False, "score": 0}
}
),
questions=[
{}
],
)For async mode, call a_create and await it as shown below:
result = await client.metrics.a_create(...)Parameters
| Parameter | Type | Description |
|---|---|---|
name | str | Required. The name of the metric, unique within your project. |
multi_turn | Optional[bool] | This is true when the metric evaluates conversations rather than single test cases. It decides which evaluationParams are valid and cannot be changed later. |
criteria | Optional[str] | The criteria the metric scores against. A GEVAL metric needs criteria or evaluationSteps. |
evaluation_steps | Optional[List[str]] | The steps the metric follows to score, as an alternative to criteria. |
evaluation_params | Optional[List[MetricEvaluationParam]] | The test case fields the metric evaluates. A single-turn metric needs at least one, and every field must match multiTurn. See MetricEvaluationParam. |
rubric | Optional[List[Rubric]] | Score ranges that anchor how the metric scores. See Rubric. |
algorithm | Optional[MetricAlgorithm] | See MetricAlgorithm. |
dag | Optional[MetricDag] | See MetricDag. |
questions | Optional[List[JevQuestion]] | The questions a JEVAL metric asks the decision model. Required when algorithm is JEVAL. See JevQuestion. |
import { ConfidentAI } from "confident-ai";
import { MetricAlgorithm } from "confident-ai/common";
const client = new ConfidentAI();
const result = await client.metrics.create(
"Correctness",
{
multiTurn: false,
criteria: "Determine if the actual output is correct based on the expected output.",
evaluationSteps: [
"Compare the actual output with the expected output.",
"Penalise any factual contradiction."
],
evaluationParams: ["actualOutput", "expectedOutput"],
rubric: [
{
scoreRange: [8, 10],
expectedOutcome: "The answer is factually correct and complete."
}
],
algorithm: MetricAlgorithm.DEFAULT,
dag: {
nodes: {
root: {
type: "BinaryJudgementNode",
criteria: "Does the actual output answer the input?",
evaluation_params: ["input", "actual_output"],
children: ["pass", "fail"]
},
pass: { type: "VerdictNode", verdict: true, score: 10 },
fail: { type: "VerdictNode", verdict: false, score: 0 }
}
},
questions: [
{}
]
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
name | string | Required. The name of the metric, unique within your project. |
multiTurn | boolean | This is true when the metric evaluates conversations rather than single test cases. It decides which evaluationParams are valid and cannot be changed later. |
criteria | string | The criteria the metric scores against. A GEVAL metric needs criteria or evaluationSteps. |
evaluationSteps | string[] | The steps the metric follows to score, as an alternative to criteria. |
evaluationParams | MetricEvaluationParam[] | The test case fields the metric evaluates. A single-turn metric needs at least one, and every field must match multiTurn. See MetricEvaluationParam. |
rubric | Rubric[] | Score ranges that anchor how the metric scores. See Rubric. |
algorithm | MetricAlgorithm | See MetricAlgorithm. |
dag | MetricDag | See MetricDag. |
questions | JevQuestion[] | The questions a JEVAL metric asks the decision model. Required when algorithm is JEVAL. See JevQuestion. |
Returns
This method returns an object of type Metric.
Get Metric
Retrieves a custom metric by id so it can be run locally. The metric must have criteria or evaluation steps and at least one evaluation parameter, or be a valid DAG.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.metrics.get(metric_id="<METRIC-ID>")For async mode, call a_get and await it as shown below:
result = await client.metrics.a_get(...)Parameters
| Parameter | Type | Description |
|---|---|---|
metric_id | str | Required. The unique id of the metric. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.metrics.get("<METRIC-ID>");Parameters
| Parameter | Type | Description |
|---|---|---|
metricId | string | Required. The unique id of the metric. |
Returns
This method returns an object of type Metric.
Update Metric
Updates a custom metric and returns it. Only the fields you send are changed; send null to clear criteria or evaluationSteps, as long as one of them remains. Every update creates a new metric version.
from confident_ai import ConfidentAI
from confident_ai.common import Rubric
client = ConfidentAI()
result = client.metrics.update(
metric_id="<METRIC-ID>",
criteria="Determine if the actual output is correct based on the expected output.",
evaluation_steps=[
"Compare the actual output with the expected output.",
"Penalise any factual contradiction."
],
evaluation_params=["actualOutput", "expectedOutput"],
rubric=[
Rubric(
score_range=[8, 10],
expected_outcome="The answer is factually correct and complete."
)
],
questions=[
{}
],
)For async mode, call a_update and await it as shown below:
result = await client.metrics.a_update(...)Parameters
| Parameter | Type | Description |
|---|---|---|
metric_id | str | Required. The unique id of the metric. |
criteria | Optional[str] | The new criteria, or null to clear it. One of criteria or evaluationSteps must remain set. |
evaluation_steps | Optional[List[str]] | The new evaluation steps, or null to clear them. One of criteria or evaluationSteps must remain set. |
evaluation_params | Optional[List[MetricEvaluationParam]] | The test case fields the metric evaluates. Each must match the metric's multiTurn. See MetricEvaluationParam. |
rubric | Optional[List[Rubric]] | Score ranges that anchor how the metric scores. See Rubric. |
questions | Optional[List[JevQuestion]] | The new questions for a JEVAL metric. Only accepted on JEVAL metrics. See JevQuestion. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.metrics.update(
"<METRIC-ID>",
{
criteria: "Determine if the actual output is correct based on the expected output.",
evaluationSteps: [
"Compare the actual output with the expected output.",
"Penalise any factual contradiction."
],
evaluationParams: ["actualOutput", "expectedOutput"],
rubric: [
{
scoreRange: [8, 10],
expectedOutcome: "The answer is factually correct and complete."
}
],
questions: [
{}
]
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
metricId | string | Required. The unique id of the metric. |
criteria | string | null | The new criteria, or null to clear it. One of criteria or evaluationSteps must remain set. |
evaluationSteps | string[] | null | The new evaluation steps, or null to clear them. One of criteria or evaluationSteps must remain set. |
evaluationParams | MetricEvaluationParam[] | The test case fields the metric evaluates. Each must match the metric's multiTurn. See MetricEvaluationParam. |
rubric | Rubric[] | Score ranges that anchor how the metric scores. See Rubric. |
questions | JevQuestion[] | The new questions for a JEVAL metric. Only accepted on JEVAL metrics. See JevQuestion. |
Returns
This method returns an object of type Metric.
Types
JevQuestion
A bounded question a JEVAL metric asks the decision model: a noul (true/false) question, a score question over ordered levels (worst first), or a choice question over options with credits.
JevQuestion = Union[JevQuestionJevQuestion0, JevQuestionJevQuestion1, JevQuestionJevQuestion2]type JevQuestion = JevQuestionJevQuestion0 | JevQuestionJevQuestion1 | JevQuestionJevQuestion2;One of .
Metric
A custom metric: how it scores, and which test case fields it needs.
class Metric:
id: str
name: str
algorithm: Optional[MetricAlgorithm]
criteria: Optional[str]
evaluation_steps: Optional[List[str]] = Field(alias="evaluationSteps")
rubric: Optional[List[Rubric]]
dag: Optional[MetricDag]
questions: Optional[List[JevQuestion]]
multi_turn: bool = Field(alias="multiTurn")
required_parameters: List[MetricEvaluationParam] = Field(alias="requiredParameters")idstrRequired
This is the unique id of the metric.
Example: "<METRIC-ID>"
namestrRequired
This is the name of the metric, unique within your project.
Example: "Correctness"
algorithmOptional[MetricAlgorithm]Required
See MetricAlgorithm.
criteriaOptional[str]Required
This is the criteria the metric scores against, or null when it uses evaluation steps.
Example: "Determine if the actual output is correct based on the expected output."
evaluation_stepsOptional[List[str]]Required
These are the steps the metric follows to score, or null when it uses criteria.
rubricOptional[List[Rubric]]Required
These are the score ranges that anchor how the metric scores, or null.
See Rubric.
dagOptional[MetricDag]Required
See MetricDag.
questionsOptional[List[JevQuestion]]Required
These are the questions a JEVAL metric asks the decision model, or null for other algorithms.
See JevQuestion.
multi_turnboolRequired
This is true when the metric evaluates conversations rather than single test cases.
Example: false
required_parametersList[MetricEvaluationParam]Required
The test case fields the metric needs to run.
Example: ["actualOutput","expectedOutput"]
interface Metric {
id: string;
name: string;
algorithm: MetricAlgorithm | null;
criteria: string | null;
evaluationSteps: string[] | null;
rubric: Rubric[] | null;
dag: MetricDag | null;
questions: JevQuestion[] | null;
multiTurn: boolean;
requiredParameters: MetricEvaluationParam[];
}idstringRequired
This is the unique id of the metric.
Example: "<METRIC-ID>"
namestringRequired
This is the name of the metric, unique within your project.
Example: "Correctness"
algorithmMetricAlgorithm | nullRequired
See MetricAlgorithm.
criteriastring | nullRequired
This is the criteria the metric scores against, or null when it uses evaluation steps.
Example: "Determine if the actual output is correct based on the expected output."
evaluationStepsstring[] | nullRequired
These are the steps the metric follows to score, or null when it uses criteria.
rubricRubric[] | nullRequired
These are the score ranges that anchor how the metric scores, or null.
See Rubric.
dagMetricDag | nullRequired
See MetricDag.
questionsJevQuestion[] | nullRequired
These are the questions a JEVAL metric asks the decision model, or null for other algorithms.
See JevQuestion.
multiTurnbooleanRequired
This is true when the metric evaluates conversations rather than single test cases.
Example: false
requiredParametersMetricEvaluationParam[]Required
The test case fields the metric needs to run.
Example: ["actualOutput","expectedOutput"]
MetricAlgorithm
The algorithm the metric is evaluated with. GEVAL scores against criteria or evaluation steps, DAG walks a decision graph, JEVAL asks a decision model bounded questions, CODE runs your own code, and DEFAULT is a built-in metric.
class MetricAlgorithm(Enum):
DEFAULT = "DEFAULT"
DAG = "DAG"
GEVAL = "GEVAL"
JEVAL = "JEVAL"
CODE = "CODE"enum MetricAlgorithm {
DEFAULT = "DEFAULT",
DAG = "DAG",
GEVAL = "GEVAL",
JEVAL = "JEVAL",
CODE = "CODE",
}DEFAULT · DAG · GEVAL · JEVAL · CODE
MetricDag
The decision graph of a DAG metric. It is validated on write, and metric references are resolved against the metrics in your project.
class MetricDag:
nodes: Dict[str, Any]nodesDict[str, Any]Required
The graph's nodes keyed by node id, as serialized by deepeval's DAG metric. Judgement nodes list their children by id; verdict nodes carry a verdict and a score, or point at another metric by metric_name.
Example: {"root":{"type":"BinaryJudgementNode","criteria":"Does the actual output answer the input?","evaluation_params":["input","actual_output"],"children":["pass","fail"]},"pass":{"type":"VerdictNode","verdict":true,"score":10},"fail":{"type":"VerdictNode","verdict":false,"score":0}}
interface MetricDag {
nodes: Record<string, unknown>;
}nodesRecord<string, unknown>Required
The graph's nodes keyed by node id, as serialized by deepeval's DAG metric. Judgement nodes list their children by id; verdict nodes carry a verdict and a score, or point at another metric by metric_name.
Example: {"root":{"type":"BinaryJudgementNode","criteria":"Does the actual output answer the input?","evaluation_params":["input","actual_output"],"children":["pass","fail"]},"pass":{"type":"VerdictNode","verdict":true,"score":10},"fail":{"type":"VerdictNode","verdict":false,"score":0}}
MetricEvaluationParam
A test case field a metric evaluates. Single-turn metrics take input, actualOutput, expectedOutput, context and expectedTools; multi-turn metrics take content, role, scenario, expectedOutcome and turns; toolsCalled, retrievalContext, metadata and tags apply to both.
class MetricEvaluationParam(Enum):
INPUT = "input"
ACTUALOUTPUT = "actualOutput"
EXPECTEDOUTPUT = "expectedOutput"
CONTEXT = "context"
EXPECTEDTOOLS = "expectedTools"
CONTENT = "content"
ROLE = "role"
SCENARIO = "scenario"
EXPECTEDOUTCOME = "expectedOutcome"
TURNS = "turns"
TOOLSCALLED = "toolsCalled"
RETRIEVALCONTEXT = "retrievalContext"
METADATA = "metadata"
TAGS = "tags"enum MetricEvaluationParam {
INPUT = "input",
ACTUALOUTPUT = "actualOutput",
EXPECTEDOUTPUT = "expectedOutput",
CONTEXT = "context",
EXPECTEDTOOLS = "expectedTools",
CONTENT = "content",
ROLE = "role",
SCENARIO = "scenario",
EXPECTEDOUTCOME = "expectedOutcome",
TURNS = "turns",
TOOLSCALLED = "toolsCalled",
RETRIEVALCONTEXT = "retrievalContext",
METADATA = "metadata",
TAGS = "tags",
}INPUT · ACTUALOUTPUT · EXPECTEDOUTPUT · CONTEXT · EXPECTEDTOOLS · CONTENT · ROLE · SCENARIO · EXPECTEDOUTCOME · TURNS · TOOLSCALLED · RETRIEVALCONTEXT · METADATA · TAGS
MetricList
class MetricList:
metrics: List[Metric]metricsList[Metric]Required
This is the list of metrics.
See Metric.
interface MetricList {
metrics: Metric[];
}metricsMetric[]Required
This is the list of metrics.
See Metric.
Rubric
A score range and the outcome it stands for. Ranges must not overlap.
class Rubric:
score_range: Tuple[int, int] = Field(alias="scoreRange")
expected_outcome: str = Field(alias="expectedOutcome")score_rangeTuple[int, int]Required
The inclusive start and end of the score range this outcome describes, each between 0 and 10, with the start no greater than the end.
Example: [8,10]
expected_outcomestrRequired
What a response scoring in this range looks like.
Example: "The answer is factually correct and complete."
interface Rubric {
scoreRange: [number, number];
expectedOutcome: string;
}scoreRange[number, number]Required
The inclusive start and end of the score range this outcome describes, each between 0 and 10, with the start no greater than the end.
Example: [8,10]
expectedOutcomestringRequired
What a response scoring in this range looks like.
Example: "The answer is factually correct and complete."
Last updated on