Metrics Batch
Overview
The Confident AI SDK exposes every Metrics Batch method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
Create Metrics Batch
Creates several GEVAL metrics at once and returns the ones created. Metrics whose name already exists in the project are skipped. DAG metrics must be created one at a time.
from confident_ai import ConfidentAI
from confident_ai.common import CreateMetricRequest
from confident_ai.common import MetricAlgorithm
from confident_ai.common import MetricDag
from confident_ai.common import Rubric
client = ConfidentAI()
result = client.metrics_batch.create(
metrics=[
CreateMetricRequest(
name="Correctness",
multi_turn=False,
criteria="Determine if the actual output is correct based on the expected output.",
evaluation_steps=[
"Compare the actual output with the expected output.",
"Penalise any factual contradiction."
],
evaluation_params=["actualOutput", "expectedOutput"],
rubric=[
Rubric(
score_range=[8, 10],
expected_outcome="The answer is factually correct and complete."
)
],
algorithm=MetricAlgorithm.DEFAULT,
dag=MetricDag(
nodes={
"root": {
"type": "BinaryJudgementNode",
"criteria": "Does the actual output answer the input?",
"evaluation_params": ["input", "actual_output"],
"children": ["pass", "fail"]
},
"pass": {"type": "VerdictNode", "verdict": True, "score": 10},
"fail": {"type": "VerdictNode", "verdict": False, "score": 0}
}
),
questions=[
{}
]
)
],
)For async mode, call a_create and await it as shown below:
result = await client.metrics_batch.a_create(...)Parameters
| Parameter | Type | Description |
|---|---|---|
metrics | List[CreateMetricRequest] | Required. The metrics to create. Names must be unique within the batch for the same multiTurn, and DAG metrics are not accepted here. See CreateMetricRequest. |
import { ConfidentAI } from "confident-ai";
import { MetricAlgorithm } from "confident-ai/common";
const client = new ConfidentAI();
const result = await client.metricsBatch.create(
[
{
name: "Correctness",
multiTurn: false,
criteria: "Determine if the actual output is correct based on the expected output.",
evaluationSteps: [
"Compare the actual output with the expected output.",
"Penalise any factual contradiction."
],
evaluationParams: ["actualOutput", "expectedOutput"],
rubric: [
{
scoreRange: [8, 10],
expectedOutcome: "The answer is factually correct and complete."
}
],
algorithm: MetricAlgorithm.DEFAULT,
dag: {
nodes: {
root: {
type: "BinaryJudgementNode",
criteria: "Does the actual output answer the input?",
evaluation_params: ["input", "actual_output"],
children: ["pass", "fail"]
},
pass: { type: "VerdictNode", verdict: true, score: 10 },
fail: { type: "VerdictNode", verdict: false, score: 0 }
}
},
questions: [
{}
]
}
],
);Parameters
| Parameter | Type | Description |
|---|---|---|
metrics | CreateMetricRequest[] | Required. The metrics to create. Names must be unique within the batch for the same multiTurn, and DAG metrics are not accepted here. See CreateMetricRequest. |
Returns
This method returns an object of type MetricList.
Types
CreateMetricRequest
A metric to create. GEVAL metrics need criteria or evaluationSteps; a DAG metric needs algorithm set to DAG and a dag; a JEVAL metric needs algorithm set to JEVAL and questions.
class CreateMetricRequest:
name: str
multi_turn: Optional[bool] = Field(default=None, alias="multiTurn")
criteria: Optional[str] = None
evaluation_steps: Optional[List[str]] = Field(default=None, alias="evaluationSteps")
evaluation_params: Optional[List[MetricEvaluationParam]] = Field(default=None, alias="evaluationParams")
rubric: Optional[List[Rubric]] = None
algorithm: Optional[MetricAlgorithm] = None
dag: Optional[MetricDag] = None
questions: Optional[List[JevQuestion]] = NonenamestrRequired
The name of the metric, unique within your project.
Example: "Correctness"
multi_turnOptional[bool]
This is true when the metric evaluates conversations rather than single test cases. It decides which evaluationParams are valid and cannot be changed later.
Example: false
criteriaOptional[str]
The criteria the metric scores against. A GEVAL metric needs criteria or evaluationSteps.
Example: "Determine if the actual output is correct based on the expected output."
evaluation_stepsOptional[List[str]]
The steps the metric follows to score, as an alternative to criteria.
Example: ["Compare the actual output with the expected output.","Penalise any factual contradiction."]
evaluation_paramsOptional[List[MetricEvaluationParam]]
The test case fields the metric evaluates. A single-turn metric needs at least one, and every field must match multiTurn.
Example: ["actualOutput","expectedOutput"]
rubricOptional[List[Rubric]]
Score ranges that anchor how the metric scores.
See Rubric.
algorithmOptional[MetricAlgorithm]
See MetricAlgorithm.
dagOptional[MetricDag]
See MetricDag.
questionsOptional[List[JevQuestion]]
The questions a JEVAL metric asks the decision model. Required when algorithm is JEVAL.
See JevQuestion.
interface CreateMetricRequest {
name: string;
multiTurn?: boolean;
criteria?: string;
evaluationSteps?: string[];
evaluationParams?: MetricEvaluationParam[];
rubric?: Rubric[];
algorithm?: MetricAlgorithm;
dag?: MetricDag;
questions?: JevQuestion[];
}namestringRequired
The name of the metric, unique within your project.
Example: "Correctness"
multiTurnboolean
This is true when the metric evaluates conversations rather than single test cases. It decides which evaluationParams are valid and cannot be changed later.
Example: false
criteriastring
The criteria the metric scores against. A GEVAL metric needs criteria or evaluationSteps.
Example: "Determine if the actual output is correct based on the expected output."
evaluationStepsstring[]
The steps the metric follows to score, as an alternative to criteria.
Example: ["Compare the actual output with the expected output.","Penalise any factual contradiction."]
evaluationParamsMetricEvaluationParam[]
The test case fields the metric evaluates. A single-turn metric needs at least one, and every field must match multiTurn.
Example: ["actualOutput","expectedOutput"]
rubricRubric[]
Score ranges that anchor how the metric scores.
See Rubric.
algorithmMetricAlgorithm
See MetricAlgorithm.
dagMetricDag
See MetricDag.
questionsJevQuestion[]
The questions a JEVAL metric asks the decision model. Required when algorithm is JEVAL.
See JevQuestion.
JevQuestion
A bounded question a JEVAL metric asks the decision model: a noul (true/false) question, a score question over ordered levels (worst first), or a choice question over options with credits.
JevQuestion = Union[JevQuestionJevQuestion0, JevQuestionJevQuestion1, JevQuestionJevQuestion2]type JevQuestion = JevQuestionJevQuestion0 | JevQuestionJevQuestion1 | JevQuestionJevQuestion2;One of .
Metric
A custom metric: how it scores, and which test case fields it needs.
class Metric:
id: str
name: str
algorithm: Optional[MetricAlgorithm]
criteria: Optional[str]
evaluation_steps: Optional[List[str]] = Field(alias="evaluationSteps")
rubric: Optional[List[Rubric]]
dag: Optional[MetricDag]
questions: Optional[List[JevQuestion]]
multi_turn: bool = Field(alias="multiTurn")
required_parameters: List[MetricEvaluationParam] = Field(alias="requiredParameters")idstrRequired
This is the unique id of the metric.
Example: "<METRIC-ID>"
namestrRequired
This is the name of the metric, unique within your project.
Example: "Correctness"
algorithmOptional[MetricAlgorithm]Required
See MetricAlgorithm.
criteriaOptional[str]Required
This is the criteria the metric scores against, or null when it uses evaluation steps.
Example: "Determine if the actual output is correct based on the expected output."
evaluation_stepsOptional[List[str]]Required
These are the steps the metric follows to score, or null when it uses criteria.
rubricOptional[List[Rubric]]Required
These are the score ranges that anchor how the metric scores, or null.
See Rubric.
dagOptional[MetricDag]Required
See MetricDag.
questionsOptional[List[JevQuestion]]Required
These are the questions a JEVAL metric asks the decision model, or null for other algorithms.
See JevQuestion.
multi_turnboolRequired
This is true when the metric evaluates conversations rather than single test cases.
Example: false
required_parametersList[MetricEvaluationParam]Required
The test case fields the metric needs to run.
Example: ["actualOutput","expectedOutput"]
interface Metric {
id: string;
name: string;
algorithm: MetricAlgorithm | null;
criteria: string | null;
evaluationSteps: string[] | null;
rubric: Rubric[] | null;
dag: MetricDag | null;
questions: JevQuestion[] | null;
multiTurn: boolean;
requiredParameters: MetricEvaluationParam[];
}idstringRequired
This is the unique id of the metric.
Example: "<METRIC-ID>"
namestringRequired
This is the name of the metric, unique within your project.
Example: "Correctness"
algorithmMetricAlgorithm | nullRequired
See MetricAlgorithm.
criteriastring | nullRequired
This is the criteria the metric scores against, or null when it uses evaluation steps.
Example: "Determine if the actual output is correct based on the expected output."
evaluationStepsstring[] | nullRequired
These are the steps the metric follows to score, or null when it uses criteria.
rubricRubric[] | nullRequired
These are the score ranges that anchor how the metric scores, or null.
See Rubric.
dagMetricDag | nullRequired
See MetricDag.
questionsJevQuestion[] | nullRequired
These are the questions a JEVAL metric asks the decision model, or null for other algorithms.
See JevQuestion.
multiTurnbooleanRequired
This is true when the metric evaluates conversations rather than single test cases.
Example: false
requiredParametersMetricEvaluationParam[]Required
The test case fields the metric needs to run.
Example: ["actualOutput","expectedOutput"]
MetricAlgorithm
The algorithm the metric is evaluated with. GEVAL scores against criteria or evaluation steps, DAG walks a decision graph, JEVAL asks a decision model bounded questions, CODE runs your own code, and DEFAULT is a built-in metric.
class MetricAlgorithm(Enum):
DEFAULT = "DEFAULT"
DAG = "DAG"
GEVAL = "GEVAL"
JEVAL = "JEVAL"
CODE = "CODE"enum MetricAlgorithm {
DEFAULT = "DEFAULT",
DAG = "DAG",
GEVAL = "GEVAL",
JEVAL = "JEVAL",
CODE = "CODE",
}DEFAULT · DAG · GEVAL · JEVAL · CODE
MetricDag
The decision graph of a DAG metric. It is validated on write, and metric references are resolved against the metrics in your project.
class MetricDag:
nodes: Dict[str, Any]nodesDict[str, Any]Required
The graph's nodes keyed by node id, as serialized by deepeval's DAG metric. Judgement nodes list their children by id; verdict nodes carry a verdict and a score, or point at another metric by metric_name.
Example: {"root":{"type":"BinaryJudgementNode","criteria":"Does the actual output answer the input?","evaluation_params":["input","actual_output"],"children":["pass","fail"]},"pass":{"type":"VerdictNode","verdict":true,"score":10},"fail":{"type":"VerdictNode","verdict":false,"score":0}}
interface MetricDag {
nodes: Record<string, unknown>;
}nodesRecord<string, unknown>Required
The graph's nodes keyed by node id, as serialized by deepeval's DAG metric. Judgement nodes list their children by id; verdict nodes carry a verdict and a score, or point at another metric by metric_name.
Example: {"root":{"type":"BinaryJudgementNode","criteria":"Does the actual output answer the input?","evaluation_params":["input","actual_output"],"children":["pass","fail"]},"pass":{"type":"VerdictNode","verdict":true,"score":10},"fail":{"type":"VerdictNode","verdict":false,"score":0}}
MetricEvaluationParam
A test case field a metric evaluates. Single-turn metrics take input, actualOutput, expectedOutput, context and expectedTools; multi-turn metrics take content, role, scenario, expectedOutcome and turns; toolsCalled, retrievalContext, metadata and tags apply to both.
class MetricEvaluationParam(Enum):
INPUT = "input"
ACTUALOUTPUT = "actualOutput"
EXPECTEDOUTPUT = "expectedOutput"
CONTEXT = "context"
EXPECTEDTOOLS = "expectedTools"
CONTENT = "content"
ROLE = "role"
SCENARIO = "scenario"
EXPECTEDOUTCOME = "expectedOutcome"
TURNS = "turns"
TOOLSCALLED = "toolsCalled"
RETRIEVALCONTEXT = "retrievalContext"
METADATA = "metadata"
TAGS = "tags"enum MetricEvaluationParam {
INPUT = "input",
ACTUALOUTPUT = "actualOutput",
EXPECTEDOUTPUT = "expectedOutput",
CONTEXT = "context",
EXPECTEDTOOLS = "expectedTools",
CONTENT = "content",
ROLE = "role",
SCENARIO = "scenario",
EXPECTEDOUTCOME = "expectedOutcome",
TURNS = "turns",
TOOLSCALLED = "toolsCalled",
RETRIEVALCONTEXT = "retrievalContext",
METADATA = "metadata",
TAGS = "tags",
}INPUT · ACTUALOUTPUT · EXPECTEDOUTPUT · CONTEXT · EXPECTEDTOOLS · CONTENT · ROLE · SCENARIO · EXPECTEDOUTCOME · TURNS · TOOLSCALLED · RETRIEVALCONTEXT · METADATA · TAGS
MetricList
class MetricList:
metrics: List[Metric]metricsList[Metric]Required
This is the list of metrics.
See Metric.
interface MetricList {
metrics: Metric[];
}metricsMetric[]Required
This is the list of metrics.
See Metric.
Rubric
A score range and the outcome it stands for. Ranges must not overlap.
class Rubric:
score_range: Tuple[int, int] = Field(alias="scoreRange")
expected_outcome: str = Field(alias="expectedOutcome")score_rangeTuple[int, int]Required
The inclusive start and end of the score range this outcome describes, each between 0 and 10, with the start no greater than the end.
Example: [8,10]
expected_outcomestrRequired
What a response scoring in this range looks like.
Example: "The answer is factually correct and complete."
interface Rubric {
scoreRange: [number, number];
expectedOutcome: string;
}scoreRange[number, number]Required
The inclusive start and end of the score range this outcome describes, each between 0 and 10, with the start no greater than the end.
Example: [8,10]
expectedOutcomestringRequired
What a response scoring in this range looks like.
Example: "The answer is factually correct and complete."
Last updated on