Launch Week 3: Five days of launches

Metrics Batch

Overview

The Confident AI SDK exposes every Metrics Batch method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

Create Metrics Batch

Creates several GEVAL metrics at once and returns the ones created. Metrics whose name already exists in the project are skipped. DAG metrics must be created one at a time.

from confident_ai import ConfidentAI
from confident_ai.common import CreateMetricRequest
from confident_ai.common import MetricAlgorithm
from confident_ai.common import MetricDag
from confident_ai.common import Rubric

client = ConfidentAI()

result = client.metrics_batch.create(
    metrics=[
        CreateMetricRequest(
            name="Correctness",
            multi_turn=False,
            criteria="Determine if the actual output is correct based on the expected output.",
            evaluation_steps=[
                "Compare the actual output with the expected output.",
                "Penalise any factual contradiction."
            ],
            evaluation_params=["actualOutput", "expectedOutput"],
            rubric=[
                Rubric(
                    score_range=[8, 10],
                    expected_outcome="The answer is factually correct and complete."
                )
            ],
            algorithm=MetricAlgorithm.DEFAULT,
            dag=MetricDag(
                nodes={
                    "root": {
                        "type": "BinaryJudgementNode",
                        "criteria": "Does the actual output answer the input?",
                        "evaluation_params": ["input", "actual_output"],
                        "children": ["pass", "fail"]
                    },
                    "pass": {"type": "VerdictNode", "verdict": True, "score": 10},
                    "fail": {"type": "VerdictNode", "verdict": False, "score": 0}
                }
            ),
            questions=[
                {}
            ]
        )
    ],
)

For async mode, call a_create and await it as shown below:

result = await client.metrics_batch.a_create(...)

Parameters

ParameterTypeDescription
metricsList[CreateMetricRequest]Required. The metrics to create. Names must be unique within the batch for the same multiTurn, and DAG metrics are not accepted here. See CreateMetricRequest.

Returns

This method returns an object of type MetricList.

Types

CreateMetricRequest

A metric to create. GEVAL metrics need criteria or evaluationSteps; a DAG metric needs algorithm set to DAG and a dag; a JEVAL metric needs algorithm set to JEVAL and questions.

class CreateMetricRequest:
    name: str
    multi_turn: Optional[bool] = Field(default=None, alias="multiTurn")
    criteria: Optional[str] = None
    evaluation_steps: Optional[List[str]] = Field(default=None, alias="evaluationSteps")
    evaluation_params: Optional[List[MetricEvaluationParam]] = Field(default=None, alias="evaluationParams")
    rubric: Optional[List[Rubric]] = None
    algorithm: Optional[MetricAlgorithm] = None
    dag: Optional[MetricDag] = None
    questions: Optional[List[JevQuestion]] = None

namestrRequired

The name of the metric, unique within your project.

Example: "Correctness"

multi_turnOptional[bool]

This is true when the metric evaluates conversations rather than single test cases. It decides which evaluationParams are valid and cannot be changed later.

Example: false

criteriaOptional[str]

The criteria the metric scores against. A GEVAL metric needs criteria or evaluationSteps.

Example: "Determine if the actual output is correct based on the expected output."

evaluation_stepsOptional[List[str]]

The steps the metric follows to score, as an alternative to criteria.

Example: ["Compare the actual output with the expected output.","Penalise any factual contradiction."]

evaluation_paramsOptional[List[MetricEvaluationParam]]

The test case fields the metric evaluates. A single-turn metric needs at least one, and every field must match multiTurn.

See MetricEvaluationParam.

Example: ["actualOutput","expectedOutput"]

rubricOptional[List[Rubric]]

Score ranges that anchor how the metric scores.

See Rubric.

algorithmOptional[MetricAlgorithm]

dagOptional[MetricDag]

questionsOptional[List[JevQuestion]]

The questions a JEVAL metric asks the decision model. Required when algorithm is JEVAL.

See JevQuestion.

JevQuestion

A bounded question a JEVAL metric asks the decision model: a noul (true/false) question, a score question over ordered levels (worst first), or a choice question over options with credits.

JevQuestion = Union[JevQuestionJevQuestion0, JevQuestionJevQuestion1, JevQuestionJevQuestion2]

One of .

Metric

A custom metric: how it scores, and which test case fields it needs.

class Metric:
    id: str
    name: str
    algorithm: Optional[MetricAlgorithm]
    criteria: Optional[str]
    evaluation_steps: Optional[List[str]] = Field(alias="evaluationSteps")
    rubric: Optional[List[Rubric]]
    dag: Optional[MetricDag]
    questions: Optional[List[JevQuestion]]
    multi_turn: bool = Field(alias="multiTurn")
    required_parameters: List[MetricEvaluationParam] = Field(alias="requiredParameters")

idstrRequired

This is the unique id of the metric.

Example: "<METRIC-ID>"

namestrRequired

This is the name of the metric, unique within your project.

Example: "Correctness"

algorithmOptional[MetricAlgorithm]Required

criteriaOptional[str]Required

This is the criteria the metric scores against, or null when it uses evaluation steps.

Example: "Determine if the actual output is correct based on the expected output."

evaluation_stepsOptional[List[str]]Required

These are the steps the metric follows to score, or null when it uses criteria.

rubricOptional[List[Rubric]]Required

These are the score ranges that anchor how the metric scores, or null.

See Rubric.

dagOptional[MetricDag]Required

questionsOptional[List[JevQuestion]]Required

These are the questions a JEVAL metric asks the decision model, or null for other algorithms.

See JevQuestion.

multi_turnboolRequired

This is true when the metric evaluates conversations rather than single test cases.

Example: false

required_parametersList[MetricEvaluationParam]Required

The test case fields the metric needs to run.

See MetricEvaluationParam.

Example: ["actualOutput","expectedOutput"]

MetricAlgorithm

The algorithm the metric is evaluated with. GEVAL scores against criteria or evaluation steps, DAG walks a decision graph, JEVAL asks a decision model bounded questions, CODE runs your own code, and DEFAULT is a built-in metric.

class MetricAlgorithm(Enum):
    DEFAULT = "DEFAULT"
    DAG = "DAG"
    GEVAL = "GEVAL"
    JEVAL = "JEVAL"
    CODE = "CODE"

DEFAULT · DAG · GEVAL · JEVAL · CODE

MetricDag

The decision graph of a DAG metric. It is validated on write, and metric references are resolved against the metrics in your project.

class MetricDag:
    nodes: Dict[str, Any]

nodesDict[str, Any]Required

The graph's nodes keyed by node id, as serialized by deepeval's DAG metric. Judgement nodes list their children by id; verdict nodes carry a verdict and a score, or point at another metric by metric_name.

Example: {"root":{"type":"BinaryJudgementNode","criteria":"Does the actual output answer the input?","evaluation_params":["input","actual_output"],"children":["pass","fail"]},"pass":{"type":"VerdictNode","verdict":true,"score":10},"fail":{"type":"VerdictNode","verdict":false,"score":0}}

MetricEvaluationParam

A test case field a metric evaluates. Single-turn metrics take input, actualOutput, expectedOutput, context and expectedTools; multi-turn metrics take content, role, scenario, expectedOutcome and turns; toolsCalled, retrievalContext, metadata and tags apply to both.

class MetricEvaluationParam(Enum):
    INPUT = "input"
    ACTUALOUTPUT = "actualOutput"
    EXPECTEDOUTPUT = "expectedOutput"
    CONTEXT = "context"
    EXPECTEDTOOLS = "expectedTools"
    CONTENT = "content"
    ROLE = "role"
    SCENARIO = "scenario"
    EXPECTEDOUTCOME = "expectedOutcome"
    TURNS = "turns"
    TOOLSCALLED = "toolsCalled"
    RETRIEVALCONTEXT = "retrievalContext"
    METADATA = "metadata"
    TAGS = "tags"

INPUT · ACTUALOUTPUT · EXPECTEDOUTPUT · CONTEXT · EXPECTEDTOOLS · CONTENT · ROLE · SCENARIO · EXPECTEDOUTCOME · TURNS · TOOLSCALLED · RETRIEVALCONTEXT · METADATA · TAGS

MetricList

class MetricList:
    metrics: List[Metric]

metricsList[Metric]Required

This is the list of metrics.

See Metric.

Rubric

A score range and the outcome it stands for. Ranges must not overlap.

class Rubric:
    score_range: Tuple[int, int] = Field(alias="scoreRange")
    expected_outcome: str = Field(alias="expectedOutcome")

score_rangeTuple[int, int]Required

The inclusive start and end of the score range this outcome describes, each between 0 and 10, with the start no greater than the end.

Example: [8,10]

expected_outcomestrRequired

What a response scoring in this range looks like.

Example: "The answer is factually correct and complete."

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI