Launch Week 3: Five days of launches

Metric Collections

Overview

The Confident AI SDK exposes every Metric Collection method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Metric Collections

Lists all the metric collections in your Confident AI project, each with the metrics inside it.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.metric_collections.list()

For async mode, call a_list and await it as shown below:

result = await client.metric_collections.a_list(...)

Returns

This method returns an object of type MetricCollectionList.

Create Metric Collection

Creates a metric collection from the name and metricsSettings you specify and returns it. A metric that does not exist in the project, or does not match multiTurn, rejects the whole request.

from confident_ai import ConfidentAI
from confident_ai.metric_collections import EvalMode
from confident_ai.metric_collections import MetricRef
from confident_ai.metric_collections import MetricSettingConfig
from confident_ai.common import ModelProvider

client = ConfidentAI()

result = client.metric_collections.create(
    name="RAG Quality",
    multi_turn=False,
    metrics_settings=[
        MetricSettingConfig(
            metric=MetricRef(
                name="Answer Relevancy"
            ),
            activated=True,
            threshold=0.8,
            include_reason=True,
            strict_mode=False,
            sample_rate=1,
            evaluation_model_provider=ModelProvider.OPEN_AI,
            evaluation_model_name="gpt-4o",
            decision_model_provider=ModelProvider.OPEN_AI,
            decision_model_name="jev-latest",
            eval_mode=EvalMode.LLM
        )
    ],
    sample_rate=1,
    input_transformer_id="<TRANSFORMER-ID>",
    output_transformer_id="<OUTPUT-TRANSFORMER-ID>",
)

For async mode, call a_create and await it as shown below:

result = await client.metric_collections.a_create(...)

Parameters

ParameterTypeDescription
namestrRequired. The name of the metric collection, which must be unique within your project.
multi_turnOptional[bool]This is true if the collection is multi-turn, which contains only multi-turn metrics. It cannot be changed once the collection exists.
metrics_settingsOptional[List[MetricSettingConfig]]The metrics in the collection with their settings. Each metric must exist in your project and match multiTurn. See MetricSettingConfig.
sample_rateOptional[float]The share of eligible entities the whole collection is run against, between 0 and 1. Applied on top of each metric's own sampleRate. Defaults to 1.
input_transformer_idOptional[str]The id of a transformer that reshapes the payload before evaluation. Send null to unset it.
output_transformer_idOptional[str]The id of a transformer that reshapes the result after evaluation. Send null to unset it.

Returns

This method returns an object of type MetricCollection.

Get Metric Collection

Retrieves a metric collection with every metric inside it and the settings configured for each one.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.metric_collections.get(
    metric_collection_id="<METRIC-COLLECTION-ID>",
)

For async mode, call a_get and await it as shown below:

result = await client.metric_collections.a_get(...)

Parameters

ParameterTypeDescription
metric_collection_idstrRequired. The unique id of the metric collection.

Returns

This method returns an object of type MetricCollection.

Update Metric Collection

Updates a metric collection and returns it. Only the fields you send are changed, and supplying metricsSettings replaces the whole list. multiTurn cannot be changed.

from confident_ai import ConfidentAI
from confident_ai.metric_collections import EvalMode
from confident_ai.metric_collections import MetricRef
from confident_ai.metric_collections import MetricSettingConfig
from confident_ai.common import ModelProvider

client = ConfidentAI()

result = client.metric_collections.update(
    metric_collection_id="<METRIC-COLLECTION-ID>",
    name="RAG Quality v2",
    metrics_settings=[
        MetricSettingConfig(
            metric=MetricRef(
                name="Answer Relevancy"
            ),
            activated=True,
            threshold=0.8,
            include_reason=True,
            strict_mode=False,
            sample_rate=1,
            evaluation_model_provider=ModelProvider.OPEN_AI,
            evaluation_model_name="gpt-4o",
            decision_model_provider=ModelProvider.OPEN_AI,
            decision_model_name="jev-latest",
            eval_mode=EvalMode.LLM
        )
    ],
    sample_rate=1,
    input_transformer_id="<TRANSFORMER-ID>",
    output_transformer_id="<OUTPUT-TRANSFORMER-ID>",
)

For async mode, call a_update and await it as shown below:

result = await client.metric_collections.a_update(...)

Parameters

ParameterTypeDescription
metric_collection_idstrRequired. The unique id of the metric collection.
nameOptional[str]The new name of the metric collection, which must be unique within your project.
metrics_settingsOptional[List[MetricSettingConfig]]The settings for every metric in the collection. Supplying this field replaces the entire list, so fetch the collection first and resend each metric it should keep. See MetricSettingConfig.
sample_rateOptional[float]The share of eligible entities the whole collection is run against, between 0 and 1. Applied on top of each metric's own sampleRate. Defaults to 1.
input_transformer_idOptional[str]The id of a transformer that reshapes the payload before evaluation. Send null to unset it.
output_transformer_idOptional[str]The id of a transformer that reshapes the result after evaluation. Send null to unset it.

Returns

This method returns an object of type MetricCollection.

Delete Metric Collection

Permanently deletes a metric collection. Every evaluation rule that runs this collection is deleted with it, and test runs and scheduled tasks stop pointing at it.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.metric_collections.delete(
    metric_collection_id="<METRIC-COLLECTION-ID>",
)

For async mode, call a_delete and await it as shown below:

result = await client.metric_collections.a_delete(...)

Parameters

ParameterTypeDescription
metric_collection_idstrRequired. The unique id of the metric collection.

Returns

This method returns an object of type MetricCollectionRef.

Types

EvalMode

Who decides in an LLM-as-a-judge metric. LLM: the evaluation model runs the whole metric. HYBRID: the evaluation model extracts and explains, and the decision model answers each decision, falling back to the evaluation model if a decision fails. DECISION: the decision model runs the whole metric in one request, with no evaluation model.

class EvalMode(Enum):
    LLM = "LLM"
    HYBRID = "HYBRID"
    DECISION = "DECISION"

LLM · HYBRID · DECISION

MetricCollection

A metric collection: its name, sampling and transformer configuration, and the settings for every metric inside it.

class MetricCollection:
    id: str
    name: str
    multi_turn: bool = Field(alias="multiTurn")
    sample_rate: float = Field(alias="sampleRate")
    input_transformer_id: Optional[str] = Field(alias="inputTransformerId")
    output_transformer_id: Optional[str] = Field(alias="outputTransformerId")
    metrics_settings: List[MetricSetting] = Field(alias="metricsSettings")

idstrRequired

This is the id of the metric collection.

Example: "<METRIC-COLLECTION-ID>"

namestrRequired

This is the name of the metric collection, which you supply to the evals API to run evaluations remotely.

Example: "RAG Quality"

multi_turnboolRequired

Whether this is a multi-turn collection. Multi-turn collections contain only multi-turn metrics and evaluate conversations rather than single test cases.

Example: false

sample_ratefloatRequired

The share of eligible entities this collection is run against, between 0 and 1. Applied on top of each metric's own sampleRate.

Example: 1

input_transformer_idOptional[str]Required

The id of the transformer that reshapes the payload before evaluation, or null when the collection does not use one.

Example: "<TRANSFORMER-ID>"

output_transformer_idOptional[str]Required

The id of the transformer that reshapes the result after evaluation, or null when the collection does not use one.

metrics_settingsList[MetricSetting]Required

The metrics in the collection with their settings.

See MetricSetting.

MetricCollectionList

class MetricCollectionList:
    metric_collections: List[MetricCollection] = Field(alias="metricCollections")

metric_collectionsList[MetricCollection]Required

This is the list of metric collections in your project.

See MetricCollection.

MetricCollectionRef

class MetricCollectionRef:
    id: str

idstrRequired

This is the id of the metric collection.

Example: "<METRIC-COLLECTION-ID>"

MetricRef

A metric referenced by its name.

class MetricRef:
    name: str

namestrRequired

The name of the metric, which must match a metric in your project or one of Confident AI's built-in metrics.

Example: "Answer Relevancy"

MetricSetting

A metric in the collection with the settings that decide whether and how it runs.

class MetricSetting:
    metric: MetricSummary
    activated: bool
    threshold: float
    include_reason: bool = Field(alias="includeReason")
    strict_mode: bool = Field(alias="strictMode")
    sample_rate: float = Field(alias="sampleRate")
    evaluation_model_provider: Optional[ModelProvider] = Field(alias="evaluationModelProvider")
    evaluation_model_name: Optional[str] = Field(alias="evaluationModelName")
    decision_model_provider: Optional[ModelProvider] = Field(alias="decisionModelProvider")
    decision_model_name: Optional[str] = Field(alias="decisionModelName")
    eval_mode: Optional[EvalMode] = Field(alias="evalMode")

metricMetricSummaryRequired

activatedboolRequired

Whether this metric is activated. Only activated metrics are run during an evaluation.

Example: true

thresholdfloatRequired

The threshold this metric is scored against. A metric passes when its score is equal to or greater than the threshold.

Example: 0.8

include_reasonboolRequired

Whether a written reason explaining the metric's score is generated during evaluation.

Example: true

strict_modeboolRequired

Whether this metric runs in strict mode, which outputs a binary score of 0 or 1 instead of a continuous score.

Example: false

sample_ratefloatRequired

The probability that this metric is run for any given evaluation, between 0 and 1. Applied on top of the collection's own sampleRate.

Example: 1

evaluation_model_providerOptional[ModelProvider]Required

evaluation_model_nameOptional[str]Required

The name of the model this metric is evaluated with, or null when the project's default evaluation model is used.

Example: "gpt-4o"

decision_model_providerOptional[ModelProvider]Required

decision_model_nameOptional[str]Required

The name of the decision model this metric decides with, or null when the project's default decision model is used.

Example: "jev-latest"

eval_modeOptional[EvalMode]Required

MetricSettingConfig

A metric to include in the collection with the settings that decide whether and how it runs. Omit evaluationModelProvider to evaluate with your project's default model, and decisionModelProvider to decide with your project's default decision model.

class MetricSettingConfig:
    metric: MetricRef
    activated: Optional[bool] = None
    threshold: Optional[float] = None
    include_reason: Optional[bool] = Field(default=None, alias="includeReason")
    strict_mode: Optional[bool] = Field(default=None, alias="strictMode")
    sample_rate: Optional[float] = Field(default=None, alias="sampleRate")
    evaluation_model_provider: Optional[ModelProvider] = Field(default=None, alias="evaluationModelProvider")
    evaluation_model_name: Optional[str] = Field(default=None, alias="evaluationModelName")
    decision_model_provider: Optional[ModelProvider] = Field(default=None, alias="decisionModelProvider")
    decision_model_name: Optional[str] = Field(default=None, alias="decisionModelName")
    eval_mode: Optional[EvalMode] = Field(default=None, alias="evalMode")

metricMetricRefRequired

activatedOptional[bool]

Whether this metric is activated. Only activated metrics are run during an evaluation.

Example: true

thresholdOptional[float]

The threshold this metric is scored against. A metric passes when its score is equal to or greater than the threshold.

Example: 0.8

include_reasonOptional[bool]

Whether a written reason explaining the metric's score is generated during evaluation.

Example: true

strict_modeOptional[bool]

Whether this metric runs in strict mode, which outputs a binary score of 0 or 1 instead of a continuous score.

Example: false

sample_rateOptional[float]

The probability that this metric is run for any given evaluation, between 0 and 1. Applied on top of the collection's own sampleRate.

Example: 1

evaluation_model_providerOptional[ModelProvider]

evaluation_model_nameOptional[str]

The name of the model this metric is evaluated with. Required whenever evaluationModelProvider is set to anything other than CONFIDENT_AI, and has no effect without a provider.

Example: "gpt-4o"

decision_model_providerOptional[ModelProvider]

decision_model_nameOptional[str]

The name of the decision model this metric decides with, used in the HYBRID and DECISION eval modes. Required whenever decisionModelProvider is set, and has no effect without a provider.

Example: "jev-latest"

eval_modeOptional[EvalMode]

MetricSummary

A metric as it appears inside a collection, by id and name.

class MetricSummary:
    id: int
    name: str

idintRequired

This is the id of the metric.

Example: 1

namestrRequired

This is the name of the metric.

Example: "Answer Relevancy"

ModelProvider

This is the provider of the model.

class ModelProvider(Enum):
    OPEN_AI = "OPEN_AI"
    CUSTOM = "CUSTOM"
    CONFIDENT_AI = "CONFIDENT_AI"
    BEDROCK = "BEDROCK"
    ANTHROPIC = "ANTHROPIC"
    GEMINI = "GEMINI"
    X_AI = "X_AI"
    DEEPSEEK = "DEEPSEEK"
    MOONSHOT_AI = "MOONSHOT_AI"
    VERTEX_AI = "VERTEX_AI"
    AZURE = "AZURE"
    MISTRAL = "MISTRAL"
    PERPLEXITY = "PERPLEXITY"
    OPEN_ROUTER = "OPEN_ROUTER"
    PORTKEY = "PORTKEY"
    LITE_LLM = "LITE_LLM"
    TRUE_FOUNDRY = "TRUE_FOUNDRY"
    HUGGING_FACE = "HUGGING_FACE"
    TYPE_SAFE = "TYPE_SAFE"
    FAL = "FAL"

OPEN_AI · CUSTOM · CONFIDENT_AI · BEDROCK · ANTHROPIC · GEMINI · X_AI · DEEPSEEK · MOONSHOT_AI · VERTEX_AI · AZURE · MISTRAL · PERPLEXITY · OPEN_ROUTER · PORTKEY · LITE_LLM · TRUE_FOUNDRY · HUGGING_FACE · TYPE_SAFE · FAL

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI