Launch Week 3: Five days of launches

Metrics Data

Overview

The Confident AI SDK exposes every Metrics Data method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Metrics Data

Lists every metric result in your Confident AI project one page at a time, newest first, across all evaluations on traces, spans, threads and test cases.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.metrics_data.list(
    page=1,
    page_size=25,
    start="2025-01-01T00:00:00+00:00",
    end="2025-01-31T23:59:59+00:00",
    multi_turn="false",
    search_term="Answer Relevancy",
)

For async mode, call a_list and await it as shown below:

result = await client.metrics_data.a_list(...)

Parameters

ParameterTypeDescription
pageOptional[int]The page of metric data to return. Defaults to 1.
page_sizeOptional[int]The number of results per page, at most 100. Defaults to 25.
startOptional[str]Returns only results recorded at or after this ISO 8601 datetime.
endOptional[str]Returns only results recorded before this ISO 8601 datetime.
multi_turnOptional[Literal['true', 'false']]Filter for results evaluated on your test case type, true for multi-turn, false for single-turn. Returns both if not specified.
search_termOptional[str]Returns only results whose metric name contains this text, case-insensitively.

Returns

This method returns an object of type MetricDataList.

Types

EvaluationErrorType

Why an evaluation errored: the AI connection or a transformer failed, the evaluation model failed, the test case lacked the parameters the metric needs, or an internal error occurred.

class EvaluationErrorType(Enum):
    AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR"
    TRANSFORMER_ERROR = "TRANSFORMER_ERROR"
    EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR"
    INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS"
    INTERNAL_ERROR = "INTERNAL_ERROR"

AI_CONNECTION_ERROR · TRANSFORMER_ERROR · EVALUATION_MODEL_ERROR · INVALID_TEST_CASE_PARAMETERS · INTERNAL_ERROR

MetricData

The result of an evaluated metric, with the ids of whatever it was recorded against.

class MetricData:
    id: str
    name: str
    score: Optional[float]
    reason: Optional[str]
    success: Optional[bool]
    threshold: Optional[float]
    strict_mode: bool = Field(alias="strictMode")
    skipped: bool
    flaky: bool
    evaluation_model: Optional[str] = Field(alias="evaluationModel")
    evaluation_cost: Optional[float] = Field(alias="evaluationCost")
    error: Optional[str]
    error_type: Optional[EvaluationErrorType] = Field(alias="errorType")
    created_at: str = Field(alias="createdAt")
    evaluated_at: Optional[str] = Field(alias="evaluatedAt")
    multi_turn: bool = Field(alias="multiTurn")
    trace_uuid: Optional[str] = Field(alias="traceUuid")
    span_uuid: Optional[str] = Field(alias="spanUuid")
    thread_id: Optional[str] = Field(alias="threadId")
    test_case_id: Optional[str] = Field(alias="testCaseId")
    test_run_id: Optional[str] = Field(alias="testRunId")

idstrRequired

The unique identifier of the metric data entry.

Example: "<METRIC-DATA-ID>"

namestrRequired

The name of the metric.

Example: "Answer Relevancy"

scoreOptional[float]Required

The final metric score, or null when the metric errored or was skipped.

Example: 0.95

reasonOptional[str]Required

The reason for the metric score, generated by the evaluation model at evaluation time.

Example: "The answer directly states the capital of France."

successOptional[bool]Required

Whether the metric score is above the threshold, or null while the evaluation is still running.

Example: true

thresholdOptional[float]Required

The threshold for the metric, which determines if the metric is passing or failing.

Example: 0.5

strict_modeboolRequired

Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.

Example: false

skippedboolRequired

Whether the metric evaluation was skipped.

Example: false

flakyboolRequired

Whether the metric's verdict was non-deterministic across runs.

Example: false

evaluation_modelOptional[str]Required

The evaluation model used to run the evaluation.

Example: "gpt-4o"

evaluation_costOptional[float]Required

The cost of running the evaluation in USD.

Example: 0.0004

errorOptional[str]Required

The error message if the evaluation failed.

error_typeOptional[EvaluationErrorType]Required

created_atstrRequired

The time the metric data was created.

Example: "2025-01-15T10:30:06+00:00"

evaluated_atOptional[str]Required

The time the metric was evaluated, or null while it is still running.

Example: "2025-01-15T10:30:09+00:00"

multi_turnboolRequired

Whether this metric was evaluated on a multi-turn conversation.

Example: false

trace_uuidOptional[str]Required

The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.

Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"

span_uuidOptional[str]Required

The uuid of the span this metric was evaluated on, for component-level metrics.

thread_idOptional[str]Required

The id of the thread this metric was evaluated on, for conversation-level metrics.

test_case_idOptional[str]Required

The id of the test case this metric was evaluated on.

test_run_idOptional[str]Required

The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.

MetricDataList

One page of metric results, with the total across all pages.

class MetricDataList:
    metrics_data: List[MetricData] = Field(alias="metricsData")
    total_metrics_data: int = Field(alias="totalMetricsData")
    page: int
    page_size: int = Field(alias="pageSize")

metrics_dataList[MetricData]Required

The metric results for the current page, newest first. Each result reports whether it was evaluated on a multi-turn conversation, and carries the ids of whatever it was recorded against.

See MetricData.

total_metrics_dataintRequired

The total number of results matching the query across all pages.

Example: 120

pageintRequired

The page this response covers.

Example: 1

page_sizeintRequired

The number of results per page.

Example: 25

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI