Metrics Data
Overview
The Confident AI SDK exposes every Metrics Data method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
List Metrics Data
Lists every metric result in your Confident AI project one page at a time, newest first, across all evaluations on traces, spans, threads and test cases.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.metrics_data.list(
page=1,
page_size=25,
start="2025-01-01T00:00:00+00:00",
end="2025-01-31T23:59:59+00:00",
multi_turn="false",
search_term="Answer Relevancy",
)For async mode, call a_list and await it as shown below:
result = await client.metrics_data.a_list(...)Parameters
| Parameter | Type | Description |
|---|---|---|
page | Optional[int] | The page of metric data to return. Defaults to 1. |
page_size | Optional[int] | The number of results per page, at most 100. Defaults to 25. |
start | Optional[str] | Returns only results recorded at or after this ISO 8601 datetime. |
end | Optional[str] | Returns only results recorded before this ISO 8601 datetime. |
multi_turn | Optional[Literal['true', 'false']] | Filter for results evaluated on your test case type, true for multi-turn, false for single-turn. Returns both if not specified. |
search_term | Optional[str] | Returns only results whose metric name contains this text, case-insensitively. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.metricsData.list(
{
page: 1,
pageSize: 25,
start: "2025-01-01T00:00:00+00:00",
end: "2025-01-31T23:59:59+00:00",
multiTurn: "false",
searchTerm: "Answer Relevancy"
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
page | number | The page of metric data to return. Defaults to 1. |
pageSize | number | The number of results per page, at most 100. Defaults to 25. |
start | string | Returns only results recorded at or after this ISO 8601 datetime. |
end | string | Returns only results recorded before this ISO 8601 datetime. |
multiTurn | "true" | "false" | Filter for results evaluated on your test case type, true for multi-turn, false for single-turn. Returns both if not specified. |
searchTerm | string | Returns only results whose metric name contains this text, case-insensitively. |
Returns
This method returns an object of type MetricDataList.
Types
EvaluationErrorType
Why an evaluation errored: the AI connection or a transformer failed, the evaluation model failed, the test case lacked the parameters the metric needs, or an internal error occurred.
class EvaluationErrorType(Enum):
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR"
TRANSFORMER_ERROR = "TRANSFORMER_ERROR"
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR"
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS"
INTERNAL_ERROR = "INTERNAL_ERROR"enum EvaluationErrorType {
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR",
TRANSFORMER_ERROR = "TRANSFORMER_ERROR",
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR",
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS",
INTERNAL_ERROR = "INTERNAL_ERROR",
}AI_CONNECTION_ERROR · TRANSFORMER_ERROR · EVALUATION_MODEL_ERROR · INVALID_TEST_CASE_PARAMETERS · INTERNAL_ERROR
MetricData
The result of an evaluated metric, with the ids of whatever it was recorded against.
class MetricData:
id: str
name: str
score: Optional[float]
reason: Optional[str]
success: Optional[bool]
threshold: Optional[float]
strict_mode: bool = Field(alias="strictMode")
skipped: bool
flaky: bool
evaluation_model: Optional[str] = Field(alias="evaluationModel")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
error: Optional[str]
error_type: Optional[EvaluationErrorType] = Field(alias="errorType")
created_at: str = Field(alias="createdAt")
evaluated_at: Optional[str] = Field(alias="evaluatedAt")
multi_turn: bool = Field(alias="multiTurn")
trace_uuid: Optional[str] = Field(alias="traceUuid")
span_uuid: Optional[str] = Field(alias="spanUuid")
thread_id: Optional[str] = Field(alias="threadId")
test_case_id: Optional[str] = Field(alias="testCaseId")
test_run_id: Optional[str] = Field(alias="testRunId")idstrRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestrRequired
The name of the metric.
Example: "Answer Relevancy"
scoreOptional[float]Required
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonOptional[str]Required
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successOptional[bool]Required
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdOptional[float]Required
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strict_modeboolRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedboolRequired
Whether the metric evaluation was skipped.
Example: false
flakyboolRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluation_modelOptional[str]Required
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluation_costOptional[float]Required
The cost of running the evaluation in USD.
Example: 0.0004
errorOptional[str]Required
The error message if the evaluation failed.
error_typeOptional[EvaluationErrorType]Required
See EvaluationErrorType.
created_atstrRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluated_atOptional[str]Required
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
multi_turnboolRequired
Whether this metric was evaluated on a multi-turn conversation.
Example: false
trace_uuidOptional[str]Required
The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
span_uuidOptional[str]Required
The uuid of the span this metric was evaluated on, for component-level metrics.
thread_idOptional[str]Required
The id of the thread this metric was evaluated on, for conversation-level metrics.
test_case_idOptional[str]Required
The id of the test case this metric was evaluated on.
test_run_idOptional[str]Required
The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.
interface MetricData {
id: string;
name: string;
score: number | null;
reason: string | null;
success: boolean | null;
threshold: number | null;
strictMode: boolean;
skipped: boolean;
flaky: boolean;
evaluationModel: string | null;
evaluationCost: number | null;
error: string | null;
errorType: EvaluationErrorType | null;
createdAt: string;
evaluatedAt: string | null;
multiTurn: boolean;
traceUuid: string | null;
spanUuid: string | null;
threadId: string | null;
testCaseId: string | null;
testRunId: string | null;
}idstringRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestringRequired
The name of the metric.
Example: "Answer Relevancy"
scorenumber | nullRequired
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonstring | nullRequired
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successboolean | nullRequired
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdnumber | nullRequired
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strictModebooleanRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedbooleanRequired
Whether the metric evaluation was skipped.
Example: false
flakybooleanRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluationModelstring | nullRequired
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluationCostnumber | nullRequired
The cost of running the evaluation in USD.
Example: 0.0004
errorstring | nullRequired
The error message if the evaluation failed.
errorTypeEvaluationErrorType | nullRequired
See EvaluationErrorType.
createdAtstringRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluatedAtstring | nullRequired
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
multiTurnbooleanRequired
Whether this metric was evaluated on a multi-turn conversation.
Example: false
traceUuidstring | nullRequired
The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
spanUuidstring | nullRequired
The uuid of the span this metric was evaluated on, for component-level metrics.
threadIdstring | nullRequired
The id of the thread this metric was evaluated on, for conversation-level metrics.
testCaseIdstring | nullRequired
The id of the test case this metric was evaluated on.
testRunIdstring | nullRequired
The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.
MetricDataList
One page of metric results, with the total across all pages.
class MetricDataList:
metrics_data: List[MetricData] = Field(alias="metricsData")
total_metrics_data: int = Field(alias="totalMetricsData")
page: int
page_size: int = Field(alias="pageSize")metrics_dataList[MetricData]Required
The metric results for the current page, newest first. Each result reports whether it was evaluated on a multi-turn conversation, and carries the ids of whatever it was recorded against.
See MetricData.
total_metrics_dataintRequired
The total number of results matching the query across all pages.
Example: 120
pageintRequired
The page this response covers.
Example: 1
page_sizeintRequired
The number of results per page.
Example: 25
interface MetricDataList {
metricsData: MetricData[];
totalMetricsData: number;
page: number;
pageSize: number;
}metricsDataMetricData[]Required
The metric results for the current page, newest first. Each result reports whether it was evaluated on a multi-turn conversation, and carries the ids of whatever it was recorded against.
See MetricData.
totalMetricsDatanumberRequired
The total number of results matching the query across all pages.
Example: 120
pagenumberRequired
The page this response covers.
Example: 1
pageSizenumberRequired
The number of results per page.
Example: 25
Last updated on