Test Runs
Overview
The Confident AI SDK exposes every Test Run method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
List Test Runs
Lists the test runs in your Confident AI project, newest first by default. Filter by status, multiTurn and the start/end window, sort with sortBy and ascending, and page with page and pageSize. Requires an active trial or paid plan.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.test_runs.list(
page=1,
page_size=25,
start="2025-01-01T00:00:00+00:00",
end="2025-01-31T23:59:59+00:00",
sort_by="createdAt",
ascending=False,
status="COMPLETED",
multi_turn=False,
)For async mode, call a_list and await it as shown below:
result = await client.test_runs.a_list(...)Parameters
| Parameter | Type | Description |
|---|---|---|
page | Optional[int] | The page of test runs to return. Defaults to 1. |
page_size | Optional[int] | The number of test runs per page, at most 100. Defaults to 25. |
start | Optional[str] | Returns only test runs created at or after this ISO 8601 datetime. Defaults to 60 days ago. |
end | Optional[str] | Returns only test runs created at or before this ISO 8601 datetime. Defaults to now. |
sort_by | Optional[Literal['createdAt', 'runDuration']] | The field to sort by. Defaults to createdAt. |
ascending | Optional[bool] | This determines if the field specified in sortBy should be in ascending order. Defaults to false. |
status | Optional[Literal['IN_PROGRESS', 'COMPLETED', 'ERRORED', 'CANCELLED']] | Returns only test runs with this status. |
multi_turn | Optional[bool] | When true, returns only multi-turn test runs; when false, only single-turn test runs. Omit to return both. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.testRuns.list(
{
page: 1,
pageSize: 25,
start: "2025-01-01T00:00:00+00:00",
end: "2025-01-31T23:59:59+00:00",
sortBy: "createdAt",
ascending: false,
status: "COMPLETED",
multiTurn: false
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
page | number | The page of test runs to return. Defaults to 1. |
pageSize | number | The number of test runs per page, at most 100. Defaults to 25. |
start | string | Returns only test runs created at or after this ISO 8601 datetime. Defaults to 60 days ago. |
end | string | Returns only test runs created at or before this ISO 8601 datetime. Defaults to now. |
sortBy | "createdAt" | "runDuration" | The field to sort by. Defaults to createdAt. |
ascending | boolean | This determines if the field specified in sortBy should be in ascending order. Defaults to false. |
status | "IN_PROGRESS" | "COMPLETED" | "ERRORED" | "CANCELLED" | Returns only test runs with this status. |
multiTurn | boolean | When true, returns only multi-turn test runs; when false, only single-turn test runs. Omit to return both. |
Returns
This method returns an object of type TestRunList.
Create Test Run
Creates a new in-progress test run and returns its id. Use this id as the testRunId when ingesting traces so that each trace becomes one test case in this run, evaluated with metricCollection when one is given.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.test_runs.create(
metric_collection="Agent Quality",
identifier="run-399-102",
)For async mode, call a_create and await it as shown below:
result = await client.test_runs.a_create(...)Parameters
| Parameter | Type | Description |
|---|---|---|
metric_collection | Optional[str] | The name of the metric collection used to evaluate the test cases formed from traces ingested into this test run. It must be a single-turn collection. |
identifier | Optional[str] | An optional human-readable identifier for the test run, shown on the Confident AI platform. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.testRuns.create(
{ metricCollection: "Agent Quality", identifier: "run-399-102" },
);Parameters
| Parameter | Type | Description |
|---|---|---|
metricCollection | string | The name of the metric collection used to evaluate the test cases formed from traces ingested into this test run. It must be a single-turn collection. |
identifier | string | An optional human-readable identifier for the test run, shown on the Confident AI platform. |
Returns
This method returns an object of type TestRunRef.
Submit Test Case Result
Submits the result for a single test case in a long-running agent evaluation. Confident AI evaluates it and finalizes the test run once every result has been received; a repeated submission for the same test case is ignored and reported as already_received. Long-running mode is available for single-turn AI connection evaluations only. Responds 410 when the testCaseId is unknown or its result window has closed, 409 when the test run is no longer accepting results, and 400 when the test run is multi-turn or has no metric collection.
from confident_ai import ConfidentAI
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
client = ConfidentAI()
result = client.test_runs.submit_test_case_result(
test_case_id="<TEST-CASE-ID>",
actual_output="Mount Everest is 8,848 metres tall.",
retrieval_context=["Everest is 8,848 metres tall."],
tools_called=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
expected_tools=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
token_cost=0.002,
input_token_count=24,
output_token_count=12,
metadata={"region": "Nepal"},
)For async mode, call a_submit_test_case_result and await it as shown below:
result = await client.test_runs.a_submit_test_case_result(...)Parameters
| Parameter | Type | Description |
|---|---|---|
test_case_id | str | Required. The test case id Confident AI sent to your AI connection as confident.testCaseId when it dispatched this golden. |
actual_output | Optional[str] | The actual output produced by your agent. |
retrieval_context | Optional[List[str]] | The retrieval context your agent used, if any. |
tools_called | Optional[List[ToolCall]] | The tools your agent called while producing the output. See ToolCall. |
expected_tools | Optional[List[ToolCall]] | The tools you expected to be called for this test case. See ToolCall. |
token_cost | Optional[float] | This is the cost of the tokens used to produce the output. |
input_token_count | Optional[int] | This is the number of input tokens passed to the LLM model. |
output_token_count | Optional[int] | This is the number of output tokens generated by the LLM model. |
metadata | Optional[Dict[str, Any]] | Optional additional metadata to attach to the test case. |
import { ConfidentAI } from "confident-ai";
import { ToolCallType } from "confident-ai/common";
const client = new ConfidentAI();
const result = await client.testRuns.submitTestCaseResult(
"<TEST-CASE-ID>",
{
actualOutput: "Mount Everest is 8,848 metres tall.",
retrievalContext: ["Everest is 8,848 metres tall."],
toolsCalled: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
expectedTools: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
tokenCost: 0.002,
inputTokenCount: 24,
outputTokenCount: 12,
metadata: { region: "Nepal" }
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
testCaseId | string | Required. The test case id Confident AI sent to your AI connection as confident.testCaseId when it dispatched this golden. |
actualOutput | string | The actual output produced by your agent. |
retrievalContext | string[] | The retrieval context your agent used, if any. |
toolsCalled | ToolCall[] | The tools your agent called while producing the output. See ToolCall. |
expectedTools | ToolCall[] | The tools you expected to be called for this test case. See ToolCall. |
tokenCost | number | This is the cost of the tokens used to produce the output. |
inputTokenCount | number | This is the number of input tokens passed to the LLM model. |
outputTokenCount | number | This is the number of output tokens generated by the LLM model. |
metadata | Record<string, unknown> | Optional additional metadata to attach to the test case. |
Returns
This method returns an object of type SubmittedTestCaseResult.
Get Test Run
Retrieves a test run by id, with its aggregated scores and the results of each test case it covers. Requires an active trial or paid plan.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.test_runs.get(test_run_id="<TEST-RUN-ID>")For async mode, call a_get and await it as shown below:
result = await client.test_runs.a_get(...)Parameters
| Parameter | Type | Description |
|---|---|---|
test_run_id | str | Required. The id of the test run. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.testRuns.get("<TEST-RUN-ID>");Parameters
| Parameter | Type | Description |
|---|---|---|
testRunId | string | Required. The id of the test run. |
Returns
This method returns an object of type TestRun.
Types
AiInsightsSummary
The AI-generated summary of a test run, produced on the Confident AI platform. Null until it has been generated.
class AiInsightsSummary:
summary_overview: SummaryOverview = Field(alias="summaryOverview")
topic_summaries: List[TopicSummary] = Field(alias="topicSummaries")summary_overviewSummaryOverviewRequired
See SummaryOverview.
topic_summariesList[TopicSummary]Required
The findings for each topic the test cases were grouped into.
See TopicSummary.
interface AiInsightsSummary {
summaryOverview: SummaryOverview;
topicSummaries: TopicSummary[];
}summaryOverviewSummaryOverviewRequired
See SummaryOverview.
topicSummariesTopicSummary[]Required
The findings for each topic the test cases were grouped into.
See TopicSummary.
Environment
This is the environment where your trace was posted, which helps with separating and debugging traces from different environments on the Confident AI platform.
class Environment(Enum):
PRODUCTION = "production"
DEVELOPMENT = "development"
STAGING = "staging"
TESTING = "testing"enum Environment {
PRODUCTION = "production",
DEVELOPMENT = "development",
STAGING = "staging",
TESTING = "testing",
}PRODUCTION · DEVELOPMENT · STAGING · TESTING
EvaluationErrorType
Why an evaluation errored: the AI connection or a transformer failed, the evaluation model failed, the test case lacked the parameters the metric needs, or an internal error occurred.
class EvaluationErrorType(Enum):
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR"
TRANSFORMER_ERROR = "TRANSFORMER_ERROR"
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR"
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS"
INTERNAL_ERROR = "INTERNAL_ERROR"enum EvaluationErrorType {
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR",
TRANSFORMER_ERROR = "TRANSFORMER_ERROR",
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR",
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS",
INTERNAL_ERROR = "INTERNAL_ERROR",
}AI_CONNECTION_ERROR · TRANSFORMER_ERROR · EVALUATION_MODEL_ERROR · INVALID_TEST_CASE_PARAMETERS · INTERNAL_ERROR
MetricScores
A metric's aggregated results across the test cases in a test run.
class MetricScores:
metric: str
scores: List[float]
passes: int
fails: int
errors: int
error_type: Optional[EvaluationErrorType] = Field(alias="errorType")metricstrRequired
This is the name of the metric.
Example: "Answer Correctness"
scoresList[float]Required
This is an array of scores for the metric across test cases, one per test case that produced a score.
Example: [0.9,1]
passesintRequired
This is the number of times this metric passed the threshold.
Example: 8
failsintRequired
This is the number of times this metric failed to pass the threshold.
Example: 2
errorsintRequired
This is the number of times this metric errored during evaluation.
Example: 0
error_typeOptional[EvaluationErrorType]Required
See EvaluationErrorType.
interface MetricScores {
metric: string;
scores: number[];
passes: number;
fails: number;
errors: number;
errorType: EvaluationErrorType | null;
}metricstringRequired
This is the name of the metric.
Example: "Answer Correctness"
scoresnumber[]Required
This is an array of scores for the metric across test cases, one per test case that produced a score.
Example: [0.9,1]
passesnumberRequired
This is the number of times this metric passed the threshold.
Example: 8
failsnumberRequired
This is the number of times this metric failed to pass the threshold.
Example: 2
errorsnumberRequired
This is the number of times this metric errored during evaluation.
Example: 0
errorTypeEvaluationErrorType | nullRequired
See EvaluationErrorType.
SubmittedTestCaseResult
The recorded test case id and whether its result was accepted.
class SubmittedTestCaseResult:
test_case_id: str = Field(alias="testCaseId")
status: TestCaseResultStatustest_case_idstrRequired
The test case id the result was recorded for.
Example: "<TEST-CASE-ID>"
statusTestCaseResultStatusRequired
See TestCaseResultStatus.
interface SubmittedTestCaseResult {
testCaseId: string;
status: TestCaseResultStatus;
}testCaseIdstringRequired
The test case id the result was recorded for.
Example: "<TEST-CASE-ID>"
statusTestCaseResultStatusRequired
See TestCaseResultStatus.
SummaryOverview
The headline findings and action items of a test run.
class SummaryOverview:
summary: List[str]
action_items: List[str] = Field(alias="actionItems")summaryList[str]Required
The headline findings across every topic.
Example: ["8 of 10 test cases passed."]
action_itemsList[str]Required
What to change to improve the next test run.
Example: ["Add goldens for lesser-known peaks."]
interface SummaryOverview {
summary: string[];
actionItems: string[];
}summarystring[]Required
The headline findings across every topic.
Example: ["8 of 10 test cases passed."]
actionItemsstring[]Required
What to change to improve the next test run.
Example: ["Add goldens for lesser-known peaks."]
SummaryPoint
A single finding in a topic summary.
class SummaryPoint:
content: str
test_case_ids: List[str] = Field(alias="testCaseIds")
grade: Optional[float] = NonecontentstrRequired
One finding about the test cases in this topic.
Example: "Answers about mountain heights were correct and cited the retrieved context."
test_case_idsList[str]Required
The ids of the test cases this finding is drawn from.
Example: ["<TEST-CASE-ID>"]
gradeOptional[float]
How well the test cases behind this finding performed, from 0 to 1.
Example: 0.9
interface SummaryPoint {
content: string;
testCaseIds: string[];
grade?: number;
}contentstringRequired
One finding about the test cases in this topic.
Example: "Answers about mountain heights were correct and cited the retrieved context."
testCaseIdsstring[]Required
The ids of the test cases this finding is drawn from.
Example: ["<TEST-CASE-ID>"]
gradenumber
How well the test cases behind this finding performed, from 0 to 1.
Example: 0.9
TestCaseMetricData
The result of evaluating one metric against a test case.
class TestCaseMetricData:
id: str
name: str
score: Optional[float]
reason: Optional[str]
success: Optional[bool]
threshold: Optional[float]
strict_mode: bool = Field(alias="strictMode")
skipped: bool
flaky: bool
evaluation_model: Optional[str] = Field(alias="evaluationModel")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
error: Optional[str]
error_type: Optional[EvaluationErrorType] = Field(alias="errorType")
created_at: str = Field(alias="createdAt")
evaluated_at: Optional[str] = Field(alias="evaluatedAt")
trace_uuid: Optional[str] = Field(alias="traceUuid")
span_uuid: Optional[str] = Field(alias="spanUuid")idstrRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestrRequired
The name of the metric.
Example: "Answer Relevancy"
scoreOptional[float]Required
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonOptional[str]Required
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successOptional[bool]Required
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdOptional[float]Required
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strict_modeboolRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedboolRequired
Whether the metric evaluation was skipped.
Example: false
flakyboolRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluation_modelOptional[str]Required
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluation_costOptional[float]Required
The cost of running the evaluation in USD.
Example: 0.0004
errorOptional[str]Required
The error message if the evaluation failed.
error_typeOptional[EvaluationErrorType]Required
See EvaluationErrorType.
created_atstrRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluated_atOptional[str]Required
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
trace_uuidOptional[str]Required
The uuid of the trace this metric was evaluated on, for test cases formed from traces.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
span_uuidOptional[str]Required
The uuid of the span this metric was evaluated on, for component-level metrics.
interface TestCaseMetricData {
id: string;
name: string;
score: number | null;
reason: string | null;
success: boolean | null;
threshold: number | null;
strictMode: boolean;
skipped: boolean;
flaky: boolean;
evaluationModel: string | null;
evaluationCost: number | null;
error: string | null;
errorType: EvaluationErrorType | null;
createdAt: string;
evaluatedAt: string | null;
traceUuid: string | null;
spanUuid: string | null;
}idstringRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestringRequired
The name of the metric.
Example: "Answer Relevancy"
scorenumber | nullRequired
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonstring | nullRequired
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successboolean | nullRequired
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdnumber | nullRequired
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strictModebooleanRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedbooleanRequired
Whether the metric evaluation was skipped.
Example: false
flakybooleanRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluationModelstring | nullRequired
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluationCostnumber | nullRequired
The cost of running the evaluation in USD.
Example: 0.0004
errorstring | nullRequired
The error message if the evaluation failed.
errorTypeEvaluationErrorType | nullRequired
See EvaluationErrorType.
createdAtstringRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluatedAtstring | nullRequired
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
traceUuidstring | nullRequired
The uuid of the trace this metric was evaluated on, for test cases formed from traces.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
spanUuidstring | nullRequired
The uuid of the span this metric was evaluated on, for component-level metrics.
TestCaseResultStatus
accepted when the result was queued for evaluation; already_received when a result for this test case had already been recorded and this retry was ignored.
class TestCaseResultStatus(Enum):
ACCEPTED = "accepted"
ALREADY_RECEIVED = "already_received"enum TestCaseResultStatus {
ACCEPTED = "accepted",
ALREADY_RECEIVED = "already_received",
}ACCEPTED · ALREADY_RECEIVED
TestCaseTrace
The trace a component-level test case was formed from, without its spans. Fetch the trace by uuid for the full span tree.
class TestCaseTrace:
uuid: str
name: Optional[str]
input: Optional[str]
output: Optional[str]
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
environment: Environment
metadata: Optional[Dict[str, Any]]
tags: Optional[List[str]]
thread_id: Optional[str] = Field(alias="threadId")
user_id: Optional[str] = Field(alias="userId")
metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
context: Optional[List[str]]
expected_output: Optional[str] = Field(alias="expectedOutput")
tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")uuidstrRequired
This is the unique identifier of the trace.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
nameOptional[str]Required
This is the name of the trace.
Example: "geography-agent"
inputOptional[str]Required
This is the input to the trace.
Example: "How tall is Mount Everest?"
outputOptional[str]Required
This is the output of the trace.
Example: "Mount Everest is 8,848 metres tall."
start_timestrRequired
This is the time the trace started.
Example: "2025-01-01T12:00:00+00:00"
end_timestrRequired
This is the time the trace ended.
Example: "2025-01-01T12:00:01.500000+00:00"
environmentEnvironmentRequired
See Environment.
metadataOptional[Dict[str, Any]]Required
This is any additional metadata associated with the trace.
Example: {"region":"Nepal"}
tagsOptional[List[str]]Required
This is any tags associated with the trace, which helps with grouping traces and filtering them on the Confident AI platform.
Example: ["geography"]
thread_idOptional[str]Required
This is the unique identifier of the thread associated with the trace.
Example: "thread-42"
user_idOptional[str]Required
This is the unique identifier for your end user for the trace.
Example: "end-user-42"
metric_collection_nameOptional[str]Required
This is the name of the metric collection the trace was evaluated with.
Example: "Agent Quality"
retrieval_contextOptional[List[str]]Required
This is the retrieval context of your trace, which is to be used for evaluation.
Example: ["Everest is 8,848 metres tall."]
contextOptional[List[str]]Required
This is the ideal retrieval context of your trace, which is to be used for evaluation.
Example: ["Everest is 8,848 metres tall."]
expected_outputOptional[str]Required
This is the expected output of your trace, which is the ideal actual output and to be used for evaluation.
Example: "Mount Everest is 8,848 metres tall."
tools_calledOptional[List[ToolCall]]Required
This is the tools called by your trace, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]Required
This is the expected tools to be called by the trace, which is to be used for evaluation.
See ToolCall.
interface TestCaseTrace {
uuid: string;
name: string | null;
input: string | null;
output: string | null;
startTime: string;
endTime: string;
environment: Environment;
metadata: Record<string, unknown> | null;
tags: string[] | null;
threadId: string | null;
userId: string | null;
metricCollectionName: string | null;
retrievalContext: string[] | null;
context: string[] | null;
expectedOutput: string | null;
toolsCalled: ToolCall[] | null;
expectedTools: ToolCall[] | null;
}uuidstringRequired
This is the unique identifier of the trace.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
namestring | nullRequired
This is the name of the trace.
Example: "geography-agent"
inputstring | nullRequired
This is the input to the trace.
Example: "How tall is Mount Everest?"
outputstring | nullRequired
This is the output of the trace.
Example: "Mount Everest is 8,848 metres tall."
startTimestringRequired
This is the time the trace started.
Example: "2025-01-01T12:00:00+00:00"
endTimestringRequired
This is the time the trace ended.
Example: "2025-01-01T12:00:01.500000+00:00"
environmentEnvironmentRequired
See Environment.
metadataRecord<string, unknown> | nullRequired
This is any additional metadata associated with the trace.
Example: {"region":"Nepal"}
tagsstring[] | nullRequired
This is any tags associated with the trace, which helps with grouping traces and filtering them on the Confident AI platform.
Example: ["geography"]
threadIdstring | nullRequired
This is the unique identifier of the thread associated with the trace.
Example: "thread-42"
userIdstring | nullRequired
This is the unique identifier for your end user for the trace.
Example: "end-user-42"
metricCollectionNamestring | nullRequired
This is the name of the metric collection the trace was evaluated with.
Example: "Agent Quality"
retrievalContextstring[] | nullRequired
This is the retrieval context of your trace, which is to be used for evaluation.
Example: ["Everest is 8,848 metres tall."]
contextstring[] | nullRequired
This is the ideal retrieval context of your trace, which is to be used for evaluation.
Example: ["Everest is 8,848 metres tall."]
expectedOutputstring | nullRequired
This is the expected output of your trace, which is the ideal actual output and to be used for evaluation.
Example: "Mount Everest is 8,848 metres tall."
toolsCalledToolCall[] | nullRequired
This is the tools called by your trace, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[] | nullRequired
This is the expected tools to be called by the trace, which is to be used for evaluation.
See ToolCall.
TestRun
A test run with its aggregated metric scores and every test case in it.
class TestRun:
id: str
created_at: str = Field(alias="createdAt")
identifier: Optional[str]
status: TestRunStatus
multi_turn: bool = Field(alias="multiTurn")
tests_passed: int = Field(alias="testsPassed")
tests_failed: int = Field(alias="testsFailed")
total_tests: int = Field(alias="totalTests")
metrics_scores: List[MetricScores] = Field(alias="metricsScores")
run_duration: float = Field(alias="runDuration")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
dataset_alias: Optional[str] = Field(alias="datasetAlias")
test_file: Optional[str] = Field(alias="testFile")
summary: Optional[AiInsightsSummary]
test_cases: List[TestRunTestCase] = Field(alias="testCases")idstrRequired
This is the unique ID for the test run, generated by Confident AI and not to be confused with the identifier provided by the user.
Example: "<TEST-RUN-ID>"
created_atstrRequired
The time the test run was created.
Example: "2025-01-01T12:00:00+00:00"
identifierOptional[str]Required
The human-readable identifier you gave the test run, if any.
Example: "run-399-102"
statusTestRunStatusRequired
See TestRunStatus.
multi_turnboolRequired
Whether this test run contains multi-turn test cases.
Example: false
tests_passedintRequired
The number of test cases that passed.
Example: 8
tests_failedintRequired
The number of test cases that failed.
Example: 2
total_testsintRequired
The total number of test cases in this test run.
Example: 10
metrics_scoresList[MetricScores]Required
The aggregated metric scores across all test cases.
See MetricScores.
run_durationfloatRequired
The total duration of the test run in seconds.
Example: 15.2
evaluation_costOptional[float]Required
The cost of evaluating every test case in the test run.
Example: 0.254
dataset_aliasOptional[str]Required
The alias of the dataset the test run was evaluated on, if any.
Example: "geography-goldens"
test_fileOptional[str]Required
The test file the test run was started from, if any.
Example: "test_geography.py"
summaryOptional[AiInsightsSummary]Required
See AiInsightsSummary.
test_casesList[TestRunTestCase]Required
The test cases in this test run. Every test case is of the same kind: single-turn, multi-turn or trace-based.
See TestRunTestCase.
interface TestRun {
id: string;
createdAt: string;
identifier: string | null;
status: TestRunStatus;
multiTurn: boolean;
testsPassed: number;
testsFailed: number;
totalTests: number;
metricsScores: MetricScores[];
runDuration: number;
evaluationCost: number | null;
datasetAlias: string | null;
testFile: string | null;
summary: AiInsightsSummary | null;
testCases: TestRunTestCase[];
}idstringRequired
This is the unique ID for the test run, generated by Confident AI and not to be confused with the identifier provided by the user.
Example: "<TEST-RUN-ID>"
createdAtstringRequired
The time the test run was created.
Example: "2025-01-01T12:00:00+00:00"
identifierstring | nullRequired
The human-readable identifier you gave the test run, if any.
Example: "run-399-102"
statusTestRunStatusRequired
See TestRunStatus.
multiTurnbooleanRequired
Whether this test run contains multi-turn test cases.
Example: false
testsPassednumberRequired
The number of test cases that passed.
Example: 8
testsFailednumberRequired
The number of test cases that failed.
Example: 2
totalTestsnumberRequired
The total number of test cases in this test run.
Example: 10
metricsScoresMetricScores[]Required
The aggregated metric scores across all test cases.
See MetricScores.
runDurationnumberRequired
The total duration of the test run in seconds.
Example: 15.2
evaluationCostnumber | nullRequired
The cost of evaluating every test case in the test run.
Example: 0.254
datasetAliasstring | nullRequired
The alias of the dataset the test run was evaluated on, if any.
Example: "geography-goldens"
testFilestring | nullRequired
The test file the test run was started from, if any.
Example: "test_geography.py"
summaryAiInsightsSummary | nullRequired
See AiInsightsSummary.
testCasesTestRunTestCase[]Required
The test cases in this test run. Every test case is of the same kind: single-turn, multi-turn or trace-based.
See TestRunTestCase.
TestRunList
class TestRunList:
test_runs: List[TestRunSummary] = Field(alias="testRuns")
total_test_runs: int = Field(alias="totalTestRuns")
page: int
page_size: int = Field(alias="pageSize")test_runsList[TestRunSummary]Required
This is the page of test runs.
See TestRunSummary.
total_test_runsintRequired
Total number of test runs matching the filters.
Example: 113
pageintRequired
The page this response covers.
Example: 1
page_sizeintRequired
The number of test runs per page.
Example: 25
interface TestRunList {
testRuns: TestRunSummary[];
totalTestRuns: number;
page: number;
pageSize: number;
}testRunsTestRunSummary[]Required
This is the page of test runs.
See TestRunSummary.
totalTestRunsnumberRequired
Total number of test runs matching the filters.
Example: 113
pagenumberRequired
The page this response covers.
Example: 1
pageSizenumberRequired
The number of test runs per page.
Example: 25
TestRunRef
class TestRunRef:
id: stridstrRequired
This is the unique identifier of the created test run. Pass it as testRunId when ingesting traces.
Example: "<TEST-RUN-ID>"
interface TestRunRef {
id: string;
}idstringRequired
This is the unique identifier of the created test run. Pass it as testRunId when ingesting traces.
Example: "<TEST-RUN-ID>"
TestRunStatus
The status of the test run: IN_PROGRESS while test cases are still being evaluated, then COMPLETED, ERRORED or CANCELLED.
class TestRunStatus(Enum):
IN_PROGRESS = "IN_PROGRESS"
COMPLETED = "COMPLETED"
ERRORED = "ERRORED"
CANCELLED = "CANCELLED"enum TestRunStatus {
IN_PROGRESS = "IN_PROGRESS",
COMPLETED = "COMPLETED",
ERRORED = "ERRORED",
CANCELLED = "CANCELLED",
}IN_PROGRESS · COMPLETED · ERRORED · CANCELLED
TestRunSummary
A test run as it appears in a listing: its counts, aggregated metric scores and AI summary, without its test cases.
class TestRunSummary:
id: str
created_at: str = Field(alias="createdAt")
identifier: Optional[str]
status: TestRunStatus
multi_turn: bool = Field(alias="multiTurn")
tests_passed: int = Field(alias="testsPassed")
tests_failed: int = Field(alias="testsFailed")
total_tests: int = Field(alias="totalTests")
metrics_scores: List[MetricScores] = Field(alias="metricsScores")
run_duration: float = Field(alias="runDuration")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
dataset_alias: Optional[str] = Field(alias="datasetAlias")
test_file: Optional[str] = Field(alias="testFile")
summary: Optional[AiInsightsSummary]idstrRequired
This is the unique ID for the test run, generated by Confident AI and not to be confused with the identifier provided by the user.
Example: "<TEST-RUN-ID>"
created_atstrRequired
The time the test run was created.
Example: "2025-01-01T12:00:00+00:00"
identifierOptional[str]Required
The human-readable identifier you gave the test run, if any.
Example: "run-399-102"
statusTestRunStatusRequired
See TestRunStatus.
multi_turnboolRequired
Whether this test run contains multi-turn test cases.
Example: false
tests_passedintRequired
The number of test cases that passed.
Example: 8
tests_failedintRequired
The number of test cases that failed.
Example: 2
total_testsintRequired
The total number of test cases in this test run.
Example: 10
metrics_scoresList[MetricScores]Required
The aggregated metric scores across all test cases.
See MetricScores.
run_durationfloatRequired
The total duration of the test run in seconds.
Example: 15.2
evaluation_costOptional[float]Required
The cost of evaluating every test case in the test run.
Example: 0.254
dataset_aliasOptional[str]Required
The alias of the dataset the test run was evaluated on, if any.
Example: "geography-goldens"
test_fileOptional[str]Required
The test file the test run was started from, if any.
Example: "test_geography.py"
summaryOptional[AiInsightsSummary]Required
See AiInsightsSummary.
interface TestRunSummary {
id: string;
createdAt: string;
identifier: string | null;
status: TestRunStatus;
multiTurn: boolean;
testsPassed: number;
testsFailed: number;
totalTests: number;
metricsScores: MetricScores[];
runDuration: number;
evaluationCost: number | null;
datasetAlias: string | null;
testFile: string | null;
summary: AiInsightsSummary | null;
}idstringRequired
This is the unique ID for the test run, generated by Confident AI and not to be confused with the identifier provided by the user.
Example: "<TEST-RUN-ID>"
createdAtstringRequired
The time the test run was created.
Example: "2025-01-01T12:00:00+00:00"
identifierstring | nullRequired
The human-readable identifier you gave the test run, if any.
Example: "run-399-102"
statusTestRunStatusRequired
See TestRunStatus.
multiTurnbooleanRequired
Whether this test run contains multi-turn test cases.
Example: false
testsPassednumberRequired
The number of test cases that passed.
Example: 8
testsFailednumberRequired
The number of test cases that failed.
Example: 2
totalTestsnumberRequired
The total number of test cases in this test run.
Example: 10
metricsScoresMetricScores[]Required
The aggregated metric scores across all test cases.
See MetricScores.
runDurationnumberRequired
The total duration of the test run in seconds.
Example: 15.2
evaluationCostnumber | nullRequired
The cost of evaluating every test case in the test run.
Example: 0.254
datasetAliasstring | nullRequired
The alias of the dataset the test run was evaluated on, if any.
Example: "geography-goldens"
testFilestring | nullRequired
The test file the test run was started from, if any.
Example: "test_geography.py"
summaryAiInsightsSummary | nullRequired
See AiInsightsSummary.
TestRunTestCase
One test case in a test run: single-turn when it carries input, multi-turn when it carries turns, trace-based when it carries trace. Every test case in one test run is of the same kind.
TestRunTestCase = Union[
TestRunSingleTurnTestCase,
TestRunMultiTurnTestCase,
TestRunTraceTestCase,
]type TestRunTestCase =
| TestRunSingleTurnTestCase
| TestRunMultiTurnTestCase
| TestRunTraceTestCase;A TestRunTestCase is one of the shapes below. Send the fields of one of them, never a mix of both.
A test case evaluated as a single exchange with your LLM application.
class TestRunSingleTurnTestCase:
input: Optional[str]
actual_output: Optional[str] = Field(alias="actualOutput")
expected_output: Optional[str] = Field(alias="expectedOutput")
context: Optional[List[str]]
retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
id: str
name: str
success: Optional[bool]
run_duration: Optional[float] = Field(alias="runDuration")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
comments: Optional[str]
additional_metadata: Optional[Dict[str, Any]] = Field(alias="additionalMetadata")
metrics_data: List[TestCaseMetricData] = Field(alias="metricsData")inputOptional[str]Required
This is the input of the test case.
Example: "How tall is Mount Everest?"
actual_outputOptional[str]Required
This is the actual output of the test case.
Example: "Mount Everest is 8,848 metres tall."
expected_outputOptional[str]Required
This is the expected output of the test case.
Example: "Mount Everest is 8,848 metres tall."
contextOptional[List[str]]Required
This is the context of the test case.
Example: ["Everest is 8,848 metres tall."]
retrieval_contextOptional[List[str]]Required
This is the retrieval context of the test case.
Example: ["Everest is 8,848 metres tall."]
tools_calledOptional[List[ToolCall]]Required
This is the tools called of the test case.
See ToolCall.
expected_toolsOptional[List[ToolCall]]Required
This is the expected tools of the test case.
See ToolCall.
idstrRequired
This is the id of the test case generated by Confident AI.
Example: "<TEST-CASE-ID>"
namestrRequired
This is the name of the test case.
Example: "everest-height"
successOptional[bool]Required
Whether this test case passed all metric thresholds, or null while it is still being evaluated.
Example: true
run_durationOptional[float]Required
The duration of the test case evaluation in seconds.
Example: 1.2
evaluation_costOptional[float]Required
The cost of evaluating this test case.
Example: 0.001
commentsOptional[str]Required
Any comments associated with this test case.
Example: "Reviewed by the geography team."
additional_metadataOptional[Dict[str, Any]]Required
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
metrics_dataList[TestCaseMetricData]Required
The metric evaluation results for this test case.
See TestCaseMetricData.
interface TestRunSingleTurnTestCase {
input: string | null;
actualOutput: string | null;
expectedOutput: string | null;
context: string[] | null;
retrievalContext: string[] | null;
toolsCalled: ToolCall[] | null;
expectedTools: ToolCall[] | null;
id: string;
name: string;
success: boolean | null;
runDuration: number | null;
evaluationCost: number | null;
comments: string | null;
additionalMetadata: Record<string, unknown> | null;
metricsData: TestCaseMetricData[];
}inputstring | nullRequired
This is the input of the test case.
Example: "How tall is Mount Everest?"
actualOutputstring | nullRequired
This is the actual output of the test case.
Example: "Mount Everest is 8,848 metres tall."
expectedOutputstring | nullRequired
This is the expected output of the test case.
Example: "Mount Everest is 8,848 metres tall."
contextstring[] | nullRequired
This is the context of the test case.
Example: ["Everest is 8,848 metres tall."]
retrievalContextstring[] | nullRequired
This is the retrieval context of the test case.
Example: ["Everest is 8,848 metres tall."]
toolsCalledToolCall[] | nullRequired
This is the tools called of the test case.
See ToolCall.
expectedToolsToolCall[] | nullRequired
This is the expected tools of the test case.
See ToolCall.
idstringRequired
This is the id of the test case generated by Confident AI.
Example: "<TEST-CASE-ID>"
namestringRequired
This is the name of the test case.
Example: "everest-height"
successboolean | nullRequired
Whether this test case passed all metric thresholds, or null while it is still being evaluated.
Example: true
runDurationnumber | nullRequired
The duration of the test case evaluation in seconds.
Example: 1.2
evaluationCostnumber | nullRequired
The cost of evaluating this test case.
Example: 0.001
commentsstring | nullRequired
Any comments associated with this test case.
Example: "Reviewed by the geography team."
additionalMetadataRecord<string, unknown> | nullRequired
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
metricsDataTestCaseMetricData[]Required
The metric evaluation results for this test case.
See TestCaseMetricData.
A test case evaluated as a conversation with your LLM application.
class TestRunMultiTurnTestCase:
turns: List[Turn]
scenario: Optional[str]
expected_outcome: Optional[str] = Field(alias="expectedOutcome")
user_description: Optional[str] = Field(alias="userDescription")
context: Optional[List[str]]
id: str
name: str
success: Optional[bool]
run_duration: Optional[float] = Field(alias="runDuration")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
comments: Optional[str]
additional_metadata: Optional[Dict[str, Any]] = Field(alias="additionalMetadata")
metrics_data: List[TestCaseMetricData] = Field(alias="metricsData")turnsList[Turn]Required
The list of turns in the conversation.
See Turn.
Example: [{"id":"<TURN-ID>","role":"user","content":"How tall is Mount Everest?"},{"id":"<TURN-ID>","role":"assistant","content":"Mount Everest is 8,848 metres tall."}]
scenarioOptional[str]Required
A description of the conversation context.
Example: "A traveller asking about mountain heights."
expected_outcomeOptional[str]Required
The expected outcome or ideal conversation flow.
Example: "The assistant states Everest's height."
user_descriptionOptional[str]Required
A description of the user in the conversation.
Example: "A traveller planning a trek."
contextOptional[List[str]]Required
The context provided for the conversation.
Example: ["Everest is 8,848 metres tall."]
idstrRequired
This is the id of the test case generated by Confident AI.
Example: "<TEST-CASE-ID>"
namestrRequired
This is the name of the test case.
Example: "everest-height"
successOptional[bool]Required
Whether this test case passed all metric thresholds, or null while it is still being evaluated.
Example: true
run_durationOptional[float]Required
The duration of the test case evaluation in seconds.
Example: 1.2
evaluation_costOptional[float]Required
The cost of evaluating this test case.
Example: 0.001
commentsOptional[str]Required
Any comments associated with this test case.
Example: "Reviewed by the geography team."
additional_metadataOptional[Dict[str, Any]]Required
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
metrics_dataList[TestCaseMetricData]Required
The metric evaluation results for this test case.
See TestCaseMetricData.
interface TestRunMultiTurnTestCase {
turns: Turn[];
scenario: string | null;
expectedOutcome: string | null;
userDescription: string | null;
context: string[] | null;
id: string;
name: string;
success: boolean | null;
runDuration: number | null;
evaluationCost: number | null;
comments: string | null;
additionalMetadata: Record<string, unknown> | null;
metricsData: TestCaseMetricData[];
}turnsTurn[]Required
The list of turns in the conversation.
See Turn.
Example: [{"id":"<TURN-ID>","role":"user","content":"How tall is Mount Everest?"},{"id":"<TURN-ID>","role":"assistant","content":"Mount Everest is 8,848 metres tall."}]
scenariostring | nullRequired
A description of the conversation context.
Example: "A traveller asking about mountain heights."
expectedOutcomestring | nullRequired
The expected outcome or ideal conversation flow.
Example: "The assistant states Everest's height."
userDescriptionstring | nullRequired
A description of the user in the conversation.
Example: "A traveller planning a trek."
contextstring[] | nullRequired
The context provided for the conversation.
Example: ["Everest is 8,848 metres tall."]
idstringRequired
This is the id of the test case generated by Confident AI.
Example: "<TEST-CASE-ID>"
namestringRequired
This is the name of the test case.
Example: "everest-height"
successboolean | nullRequired
Whether this test case passed all metric thresholds, or null while it is still being evaluated.
Example: true
runDurationnumber | nullRequired
The duration of the test case evaluation in seconds.
Example: 1.2
evaluationCostnumber | nullRequired
The cost of evaluating this test case.
Example: 0.001
commentsstring | nullRequired
Any comments associated with this test case.
Example: "Reviewed by the geography team."
additionalMetadataRecord<string, unknown> | nullRequired
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
metricsDataTestCaseMetricData[]Required
The metric evaluation results for this test case.
See TestCaseMetricData.
A test case formed from an ingested trace and evaluated component by component. Its span-level metrics are in metricsData.
class TestRunTraceTestCase:
trace: TestCaseTrace
id: str
name: str
success: Optional[bool]
run_duration: Optional[float] = Field(alias="runDuration")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
comments: Optional[str]
additional_metadata: Optional[Dict[str, Any]] = Field(alias="additionalMetadata")
metrics_data: List[TestCaseMetricData] = Field(alias="metricsData")traceTestCaseTraceRequired
See TestCaseTrace.
idstrRequired
This is the id of the test case generated by Confident AI.
Example: "<TEST-CASE-ID>"
namestrRequired
This is the name of the test case.
Example: "everest-height"
successOptional[bool]Required
Whether this test case passed all metric thresholds, or null while it is still being evaluated.
Example: true
run_durationOptional[float]Required
The duration of the test case evaluation in seconds.
Example: 1.2
evaluation_costOptional[float]Required
The cost of evaluating this test case.
Example: 0.001
commentsOptional[str]Required
Any comments associated with this test case.
Example: "Reviewed by the geography team."
additional_metadataOptional[Dict[str, Any]]Required
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
metrics_dataList[TestCaseMetricData]Required
The metric evaluation results for this test case.
See TestCaseMetricData.
interface TestRunTraceTestCase {
trace: TestCaseTrace;
id: string;
name: string;
success: boolean | null;
runDuration: number | null;
evaluationCost: number | null;
comments: string | null;
additionalMetadata: Record<string, unknown> | null;
metricsData: TestCaseMetricData[];
}traceTestCaseTraceRequired
See TestCaseTrace.
idstringRequired
This is the id of the test case generated by Confident AI.
Example: "<TEST-CASE-ID>"
namestringRequired
This is the name of the test case.
Example: "everest-height"
successboolean | nullRequired
Whether this test case passed all metric thresholds, or null while it is still being evaluated.
Example: true
runDurationnumber | nullRequired
The duration of the test case evaluation in seconds.
Example: 1.2
evaluationCostnumber | nullRequired
The cost of evaluating this test case.
Example: 0.001
commentsstring | nullRequired
Any comments associated with this test case.
Example: "Reviewed by the geography team."
additionalMetadataRecord<string, unknown> | nullRequired
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
metricsDataTestCaseMetricData[]Required
The metric evaluation results for this test case.
See TestCaseMetricData.
ToolCall
A tool your LLM application invoked, with what it passed in and what came back.
class ToolCall:
name: str
type: Optional[ToolCallType] = None
description: Optional[str] = None
input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
output: Optional[Any] = None
reasoning: Optional[str] = NonenamestrRequired
This is the name of the tool.
Example: "get_landmark_info"
typeOptional[ToolCallType]
See ToolCallType.
descriptionOptional[str]
This is the description of the tool.
Example: "This tool gives information about a mountain."
input_parametersOptional[Dict[str, Any]]
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputOptional[Any]
This is the output of the tool.
Example: "8,848 metres"
reasoningOptional[str]
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
interface ToolCall {
name: string;
type?: ToolCallType;
description?: string;
inputParameters?: Record<string, unknown> | null;
output?: unknown;
reasoning?: string;
}namestringRequired
This is the name of the tool.
Example: "get_landmark_info"
typeToolCallType
See ToolCallType.
descriptionstring
This is the description of the tool.
Example: "This tool gives information about a mountain."
inputParametersRecord<string, unknown> | null
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputunknown
This is the output of the tool.
Example: "8,848 metres"
reasoningstring
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
ToolCallType
The type of the tool call, either a function or an MCP tool.
class ToolCallType(Enum):
FUNCTION = "FUNCTION"
MCP = "MCP"enum ToolCallType {
FUNCTION = "FUNCTION",
MCP = "MCP",
}FUNCTION · MCP
TopicSummary
The findings for one topic of test cases.
class TopicSummary:
topic: str
summary_points: List[SummaryPoint] = Field(alias="summaryPoints")
test_case_ids: List[str] = Field(alias="testCaseIds")topicstrRequired
The topic the test cases were grouped under.
Example: "Mountain heights"
summary_pointsList[SummaryPoint]Required
The findings for this topic.
See SummaryPoint.
test_case_idsList[str]Required
The ids of the test cases grouped under this topic.
Example: ["<TEST-CASE-ID>"]
interface TopicSummary {
topic: string;
summaryPoints: SummaryPoint[];
testCaseIds: string[];
}topicstringRequired
The topic the test cases were grouped under.
Example: "Mountain heights"
summaryPointsSummaryPoint[]Required
The findings for this topic.
See SummaryPoint.
testCaseIdsstring[]Required
The ids of the test cases grouped under this topic.
Example: ["<TEST-CASE-ID>"]
Turn
One message in a conversation, from either the user or the assistant, with the context and tools behind an assistant reply.
class Turn:
id: Optional[str] = None
role: TurnRole
content: str
user_id: Optional[str] = Field(default=None, alias="userId")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")idOptional[str]
The id of a turn assigned by Confident AI.
Example: "<TURN-ID>"
roleTurnRoleRequired
See TurnRole.
contentstrRequired
The message content of the turn.
Example: "How tall is Mount Everest?"
user_idOptional[str]
The user ID associated with the turn.
Example: "end-user-42"
retrieval_contextOptional[List[str]]
The contexts retrieved to generate the LLM response for this turn.
Example: ["Everest is 8,848 metres tall."]
tools_calledOptional[List[ToolCall]]
The tools called to generate the LLM response for this turn.
See ToolCall.
interface Turn {
id?: string;
role: TurnRole;
content: string;
userId?: string;
retrievalContext?: string[] | null;
toolsCalled?: ToolCall[] | null;
}idstring
The id of a turn assigned by Confident AI.
Example: "<TURN-ID>"
roleTurnRoleRequired
See TurnRole.
contentstringRequired
The message content of the turn.
Example: "How tall is Mount Everest?"
userIdstring
The user ID associated with the turn.
Example: "end-user-42"
retrievalContextstring[] | null
The contexts retrieved to generate the LLM response for this turn.
Example: ["Everest is 8,848 metres tall."]
toolsCalledToolCall[] | null
The tools called to generate the LLM response for this turn.
See ToolCall.
TurnRole
The role of the turn, either user or assistant.
class TurnRole(Enum):
USER = "user"
ASSISTANT = "assistant"enum TurnRole {
USER = "user",
ASSISTANT = "assistant",
}USER · ASSISTANT
Last updated on