Launch Week 3: Five days of launches

Test Runs

Overview

The Confident AI SDK exposes every Test Run method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Test Runs

Lists the test runs in your Confident AI project, newest first by default. Filter by status, multiTurn and the start/end window, sort with sortBy and ascending, and page with page and pageSize. Requires an active trial or paid plan.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.test_runs.list(
    page=1,
    page_size=25,
    start="2025-01-01T00:00:00+00:00",
    end="2025-01-31T23:59:59+00:00",
    sort_by="createdAt",
    ascending=False,
    status="COMPLETED",
    multi_turn=False,
)

For async mode, call a_list and await it as shown below:

result = await client.test_runs.a_list(...)

Parameters

ParameterTypeDescription
pageOptional[int]The page of test runs to return. Defaults to 1.
page_sizeOptional[int]The number of test runs per page, at most 100. Defaults to 25.
startOptional[str]Returns only test runs created at or after this ISO 8601 datetime. Defaults to 60 days ago.
endOptional[str]Returns only test runs created at or before this ISO 8601 datetime. Defaults to now.
sort_byOptional[Literal['createdAt', 'runDuration']]The field to sort by. Defaults to createdAt.
ascendingOptional[bool]This determines if the field specified in sortBy should be in ascending order. Defaults to false.
statusOptional[Literal['IN_PROGRESS', 'COMPLETED', 'ERRORED', 'CANCELLED']]Returns only test runs with this status.
multi_turnOptional[bool]When true, returns only multi-turn test runs; when false, only single-turn test runs. Omit to return both.

Returns

This method returns an object of type TestRunList.

Create Test Run

Creates a new in-progress test run and returns its id. Use this id as the testRunId when ingesting traces so that each trace becomes one test case in this run, evaluated with metricCollection when one is given.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.test_runs.create(
    metric_collection="Agent Quality",
    identifier="run-399-102",
)

For async mode, call a_create and await it as shown below:

result = await client.test_runs.a_create(...)

Parameters

ParameterTypeDescription
metric_collectionOptional[str]The name of the metric collection used to evaluate the test cases formed from traces ingested into this test run. It must be a single-turn collection.
identifierOptional[str]An optional human-readable identifier for the test run, shown on the Confident AI platform.

Returns

This method returns an object of type TestRunRef.

Submit Test Case Result

Submits the result for a single test case in a long-running agent evaluation. Confident AI evaluates it and finalizes the test run once every result has been received; a repeated submission for the same test case is ignored and reported as already_received. Long-running mode is available for single-turn AI connection evaluations only. Responds 410 when the testCaseId is unknown or its result window has closed, 409 when the test run is no longer accepting results, and 400 when the test run is multi-turn or has no metric collection.

from confident_ai import ConfidentAI
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

result = client.test_runs.submit_test_case_result(
    test_case_id="<TEST-CASE-ID>",
    actual_output="Mount Everest is 8,848 metres tall.",
    retrieval_context=["Everest is 8,848 metres tall."],
    tools_called=[
        ToolCall(
            name="get_landmark_info",
            type=ToolCallType.FUNCTION,
            description="This tool gives information about a mountain.",
            input_parameters={"mountain": "Everest"},
            output="8,848 metres",
            reasoning="The user asked for the height of a mountain."
        )
    ],
    expected_tools=[
        ToolCall(
            name="get_landmark_info",
            type=ToolCallType.FUNCTION,
            description="This tool gives information about a mountain.",
            input_parameters={"mountain": "Everest"},
            output="8,848 metres",
            reasoning="The user asked for the height of a mountain."
        )
    ],
    token_cost=0.002,
    input_token_count=24,
    output_token_count=12,
    metadata={"region": "Nepal"},
)

For async mode, call a_submit_test_case_result and await it as shown below:

result = await client.test_runs.a_submit_test_case_result(...)

Parameters

ParameterTypeDescription
test_case_idstrRequired. The test case id Confident AI sent to your AI connection as confident.testCaseId when it dispatched this golden.
actual_outputOptional[str]The actual output produced by your agent.
retrieval_contextOptional[List[str]]The retrieval context your agent used, if any.
tools_calledOptional[List[ToolCall]]The tools your agent called while producing the output. See ToolCall.
expected_toolsOptional[List[ToolCall]]The tools you expected to be called for this test case. See ToolCall.
token_costOptional[float]This is the cost of the tokens used to produce the output.
input_token_countOptional[int]This is the number of input tokens passed to the LLM model.
output_token_countOptional[int]This is the number of output tokens generated by the LLM model.
metadataOptional[Dict[str, Any]]Optional additional metadata to attach to the test case.

Returns

This method returns an object of type SubmittedTestCaseResult.

Get Test Run

Retrieves a test run by id, with its aggregated scores and the results of each test case it covers. Requires an active trial or paid plan.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.test_runs.get(test_run_id="<TEST-RUN-ID>")

For async mode, call a_get and await it as shown below:

result = await client.test_runs.a_get(...)

Parameters

ParameterTypeDescription
test_run_idstrRequired. The id of the test run.

Returns

This method returns an object of type TestRun.

Types

AiInsightsSummary

The AI-generated summary of a test run, produced on the Confident AI platform. Null until it has been generated.

class AiInsightsSummary:
    summary_overview: SummaryOverview = Field(alias="summaryOverview")
    topic_summaries: List[TopicSummary] = Field(alias="topicSummaries")

summary_overviewSummaryOverviewRequired

topic_summariesList[TopicSummary]Required

The findings for each topic the test cases were grouped into.

See TopicSummary.

Environment

This is the environment where your trace was posted, which helps with separating and debugging traces from different environments on the Confident AI platform.

class Environment(Enum):
    PRODUCTION = "production"
    DEVELOPMENT = "development"
    STAGING = "staging"
    TESTING = "testing"

PRODUCTION · DEVELOPMENT · STAGING · TESTING

EvaluationErrorType

Why an evaluation errored: the AI connection or a transformer failed, the evaluation model failed, the test case lacked the parameters the metric needs, or an internal error occurred.

class EvaluationErrorType(Enum):
    AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR"
    TRANSFORMER_ERROR = "TRANSFORMER_ERROR"
    EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR"
    INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS"
    INTERNAL_ERROR = "INTERNAL_ERROR"

AI_CONNECTION_ERROR · TRANSFORMER_ERROR · EVALUATION_MODEL_ERROR · INVALID_TEST_CASE_PARAMETERS · INTERNAL_ERROR

MetricScores

A metric's aggregated results across the test cases in a test run.

class MetricScores:
    metric: str
    scores: List[float]
    passes: int
    fails: int
    errors: int
    error_type: Optional[EvaluationErrorType] = Field(alias="errorType")

metricstrRequired

This is the name of the metric.

Example: "Answer Correctness"

scoresList[float]Required

This is an array of scores for the metric across test cases, one per test case that produced a score.

Example: [0.9,1]

passesintRequired

This is the number of times this metric passed the threshold.

Example: 8

failsintRequired

This is the number of times this metric failed to pass the threshold.

Example: 2

errorsintRequired

This is the number of times this metric errored during evaluation.

Example: 0

error_typeOptional[EvaluationErrorType]Required

SubmittedTestCaseResult

The recorded test case id and whether its result was accepted.

class SubmittedTestCaseResult:
    test_case_id: str = Field(alias="testCaseId")
    status: TestCaseResultStatus

test_case_idstrRequired

The test case id the result was recorded for.

Example: "<TEST-CASE-ID>"

statusTestCaseResultStatusRequired

SummaryOverview

The headline findings and action items of a test run.

class SummaryOverview:
    summary: List[str]
    action_items: List[str] = Field(alias="actionItems")

summaryList[str]Required

The headline findings across every topic.

Example: ["8 of 10 test cases passed."]

action_itemsList[str]Required

What to change to improve the next test run.

Example: ["Add goldens for lesser-known peaks."]

SummaryPoint

A single finding in a topic summary.

class SummaryPoint:
    content: str
    test_case_ids: List[str] = Field(alias="testCaseIds")
    grade: Optional[float] = None

contentstrRequired

One finding about the test cases in this topic.

Example: "Answers about mountain heights were correct and cited the retrieved context."

test_case_idsList[str]Required

The ids of the test cases this finding is drawn from.

Example: ["<TEST-CASE-ID>"]

gradeOptional[float]

How well the test cases behind this finding performed, from 0 to 1.

Example: 0.9

TestCaseMetricData

The result of evaluating one metric against a test case.

class TestCaseMetricData:
    id: str
    name: str
    score: Optional[float]
    reason: Optional[str]
    success: Optional[bool]
    threshold: Optional[float]
    strict_mode: bool = Field(alias="strictMode")
    skipped: bool
    flaky: bool
    evaluation_model: Optional[str] = Field(alias="evaluationModel")
    evaluation_cost: Optional[float] = Field(alias="evaluationCost")
    error: Optional[str]
    error_type: Optional[EvaluationErrorType] = Field(alias="errorType")
    created_at: str = Field(alias="createdAt")
    evaluated_at: Optional[str] = Field(alias="evaluatedAt")
    trace_uuid: Optional[str] = Field(alias="traceUuid")
    span_uuid: Optional[str] = Field(alias="spanUuid")

idstrRequired

The unique identifier of the metric data entry.

Example: "<METRIC-DATA-ID>"

namestrRequired

The name of the metric.

Example: "Answer Relevancy"

scoreOptional[float]Required

The final metric score, or null when the metric errored or was skipped.

Example: 0.95

reasonOptional[str]Required

The reason for the metric score, generated by the evaluation model at evaluation time.

Example: "The answer directly states the capital of France."

successOptional[bool]Required

Whether the metric score is above the threshold, or null while the evaluation is still running.

Example: true

thresholdOptional[float]Required

The threshold for the metric, which determines if the metric is passing or failing.

Example: 0.5

strict_modeboolRequired

Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.

Example: false

skippedboolRequired

Whether the metric evaluation was skipped.

Example: false

flakyboolRequired

Whether the metric's verdict was non-deterministic across runs.

Example: false

evaluation_modelOptional[str]Required

The evaluation model used to run the evaluation.

Example: "gpt-4o"

evaluation_costOptional[float]Required

The cost of running the evaluation in USD.

Example: 0.0004

errorOptional[str]Required

The error message if the evaluation failed.

error_typeOptional[EvaluationErrorType]Required

created_atstrRequired

The time the metric data was created.

Example: "2025-01-15T10:30:06+00:00"

evaluated_atOptional[str]Required

The time the metric was evaluated, or null while it is still running.

Example: "2025-01-15T10:30:09+00:00"

trace_uuidOptional[str]Required

The uuid of the trace this metric was evaluated on, for test cases formed from traces.

Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"

span_uuidOptional[str]Required

The uuid of the span this metric was evaluated on, for component-level metrics.

TestCaseResultStatus

accepted when the result was queued for evaluation; already_received when a result for this test case had already been recorded and this retry was ignored.

class TestCaseResultStatus(Enum):
    ACCEPTED = "accepted"
    ALREADY_RECEIVED = "already_received"

ACCEPTED · ALREADY_RECEIVED

TestCaseTrace

The trace a component-level test case was formed from, without its spans. Fetch the trace by uuid for the full span tree.

class TestCaseTrace:
    uuid: str
    name: Optional[str]
    input: Optional[str]
    output: Optional[str]
    start_time: str = Field(alias="startTime")
    end_time: str = Field(alias="endTime")
    environment: Environment
    metadata: Optional[Dict[str, Any]]
    tags: Optional[List[str]]
    thread_id: Optional[str] = Field(alias="threadId")
    user_id: Optional[str] = Field(alias="userId")
    metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
    retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
    context: Optional[List[str]]
    expected_output: Optional[str] = Field(alias="expectedOutput")
    tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")

uuidstrRequired

This is the unique identifier of the trace.

Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"

nameOptional[str]Required

This is the name of the trace.

Example: "geography-agent"

inputOptional[str]Required

This is the input to the trace.

Example: "How tall is Mount Everest?"

outputOptional[str]Required

This is the output of the trace.

Example: "Mount Everest is 8,848 metres tall."

start_timestrRequired

This is the time the trace started.

Example: "2025-01-01T12:00:00+00:00"

end_timestrRequired

This is the time the trace ended.

Example: "2025-01-01T12:00:01.500000+00:00"

environmentEnvironmentRequired

metadataOptional[Dict[str, Any]]Required

This is any additional metadata associated with the trace.

Example: {"region":"Nepal"}

tagsOptional[List[str]]Required

This is any tags associated with the trace, which helps with grouping traces and filtering them on the Confident AI platform.

Example: ["geography"]

thread_idOptional[str]Required

This is the unique identifier of the thread associated with the trace.

Example: "thread-42"

user_idOptional[str]Required

This is the unique identifier for your end user for the trace.

Example: "end-user-42"

metric_collection_nameOptional[str]Required

This is the name of the metric collection the trace was evaluated with.

Example: "Agent Quality"

retrieval_contextOptional[List[str]]Required

This is the retrieval context of your trace, which is to be used for evaluation.

Example: ["Everest is 8,848 metres tall."]

contextOptional[List[str]]Required

This is the ideal retrieval context of your trace, which is to be used for evaluation.

Example: ["Everest is 8,848 metres tall."]

expected_outputOptional[str]Required

This is the expected output of your trace, which is the ideal actual output and to be used for evaluation.

Example: "Mount Everest is 8,848 metres tall."

tools_calledOptional[List[ToolCall]]Required

This is the tools called by your trace, which is to be used for evaluation.

See ToolCall.

expected_toolsOptional[List[ToolCall]]Required

This is the expected tools to be called by the trace, which is to be used for evaluation.

See ToolCall.

TestRun

A test run with its aggregated metric scores and every test case in it.

class TestRun:
    id: str
    created_at: str = Field(alias="createdAt")
    identifier: Optional[str]
    status: TestRunStatus
    multi_turn: bool = Field(alias="multiTurn")
    tests_passed: int = Field(alias="testsPassed")
    tests_failed: int = Field(alias="testsFailed")
    total_tests: int = Field(alias="totalTests")
    metrics_scores: List[MetricScores] = Field(alias="metricsScores")
    run_duration: float = Field(alias="runDuration")
    evaluation_cost: Optional[float] = Field(alias="evaluationCost")
    dataset_alias: Optional[str] = Field(alias="datasetAlias")
    test_file: Optional[str] = Field(alias="testFile")
    summary: Optional[AiInsightsSummary]
    test_cases: List[TestRunTestCase] = Field(alias="testCases")

idstrRequired

This is the unique ID for the test run, generated by Confident AI and not to be confused with the identifier provided by the user.

Example: "<TEST-RUN-ID>"

created_atstrRequired

The time the test run was created.

Example: "2025-01-01T12:00:00+00:00"

identifierOptional[str]Required

The human-readable identifier you gave the test run, if any.

Example: "run-399-102"

statusTestRunStatusRequired

multi_turnboolRequired

Whether this test run contains multi-turn test cases.

Example: false

tests_passedintRequired

The number of test cases that passed.

Example: 8

tests_failedintRequired

The number of test cases that failed.

Example: 2

total_testsintRequired

The total number of test cases in this test run.

Example: 10

metrics_scoresList[MetricScores]Required

The aggregated metric scores across all test cases.

See MetricScores.

run_durationfloatRequired

The total duration of the test run in seconds.

Example: 15.2

evaluation_costOptional[float]Required

The cost of evaluating every test case in the test run.

Example: 0.254

dataset_aliasOptional[str]Required

The alias of the dataset the test run was evaluated on, if any.

Example: "geography-goldens"

test_fileOptional[str]Required

The test file the test run was started from, if any.

Example: "test_geography.py"

summaryOptional[AiInsightsSummary]Required

test_casesList[TestRunTestCase]Required

The test cases in this test run. Every test case is of the same kind: single-turn, multi-turn or trace-based.

See TestRunTestCase.

TestRunList

class TestRunList:
    test_runs: List[TestRunSummary] = Field(alias="testRuns")
    total_test_runs: int = Field(alias="totalTestRuns")
    page: int
    page_size: int = Field(alias="pageSize")

test_runsList[TestRunSummary]Required

This is the page of test runs.

See TestRunSummary.

total_test_runsintRequired

Total number of test runs matching the filters.

Example: 113

pageintRequired

The page this response covers.

Example: 1

page_sizeintRequired

The number of test runs per page.

Example: 25

TestRunRef

class TestRunRef:
    id: str

idstrRequired

This is the unique identifier of the created test run. Pass it as testRunId when ingesting traces.

Example: "<TEST-RUN-ID>"

TestRunStatus

The status of the test run: IN_PROGRESS while test cases are still being evaluated, then COMPLETED, ERRORED or CANCELLED.

class TestRunStatus(Enum):
    IN_PROGRESS = "IN_PROGRESS"
    COMPLETED = "COMPLETED"
    ERRORED = "ERRORED"
    CANCELLED = "CANCELLED"

IN_PROGRESS · COMPLETED · ERRORED · CANCELLED

TestRunSummary

A test run as it appears in a listing: its counts, aggregated metric scores and AI summary, without its test cases.

class TestRunSummary:
    id: str
    created_at: str = Field(alias="createdAt")
    identifier: Optional[str]
    status: TestRunStatus
    multi_turn: bool = Field(alias="multiTurn")
    tests_passed: int = Field(alias="testsPassed")
    tests_failed: int = Field(alias="testsFailed")
    total_tests: int = Field(alias="totalTests")
    metrics_scores: List[MetricScores] = Field(alias="metricsScores")
    run_duration: float = Field(alias="runDuration")
    evaluation_cost: Optional[float] = Field(alias="evaluationCost")
    dataset_alias: Optional[str] = Field(alias="datasetAlias")
    test_file: Optional[str] = Field(alias="testFile")
    summary: Optional[AiInsightsSummary]

idstrRequired

This is the unique ID for the test run, generated by Confident AI and not to be confused with the identifier provided by the user.

Example: "<TEST-RUN-ID>"

created_atstrRequired

The time the test run was created.

Example: "2025-01-01T12:00:00+00:00"

identifierOptional[str]Required

The human-readable identifier you gave the test run, if any.

Example: "run-399-102"

statusTestRunStatusRequired

multi_turnboolRequired

Whether this test run contains multi-turn test cases.

Example: false

tests_passedintRequired

The number of test cases that passed.

Example: 8

tests_failedintRequired

The number of test cases that failed.

Example: 2

total_testsintRequired

The total number of test cases in this test run.

Example: 10

metrics_scoresList[MetricScores]Required

The aggregated metric scores across all test cases.

See MetricScores.

run_durationfloatRequired

The total duration of the test run in seconds.

Example: 15.2

evaluation_costOptional[float]Required

The cost of evaluating every test case in the test run.

Example: 0.254

dataset_aliasOptional[str]Required

The alias of the dataset the test run was evaluated on, if any.

Example: "geography-goldens"

test_fileOptional[str]Required

The test file the test run was started from, if any.

Example: "test_geography.py"

summaryOptional[AiInsightsSummary]Required

TestRunTestCase

One test case in a test run: single-turn when it carries input, multi-turn when it carries turns, trace-based when it carries trace. Every test case in one test run is of the same kind.

TestRunTestCase = Union[
    TestRunSingleTurnTestCase,
    TestRunMultiTurnTestCase,
    TestRunTraceTestCase,
]

A TestRunTestCase is one of the shapes below. Send the fields of one of them, never a mix of both.

A test case evaluated as a single exchange with your LLM application.

class TestRunSingleTurnTestCase:
    input: Optional[str]
    actual_output: Optional[str] = Field(alias="actualOutput")
    expected_output: Optional[str] = Field(alias="expectedOutput")
    context: Optional[List[str]]
    retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
    id: str
    name: str
    success: Optional[bool]
    run_duration: Optional[float] = Field(alias="runDuration")
    evaluation_cost: Optional[float] = Field(alias="evaluationCost")
    comments: Optional[str]
    additional_metadata: Optional[Dict[str, Any]] = Field(alias="additionalMetadata")
    metrics_data: List[TestCaseMetricData] = Field(alias="metricsData")

inputOptional[str]Required

This is the input of the test case.

Example: "How tall is Mount Everest?"

actual_outputOptional[str]Required

This is the actual output of the test case.

Example: "Mount Everest is 8,848 metres tall."

expected_outputOptional[str]Required

This is the expected output of the test case.

Example: "Mount Everest is 8,848 metres tall."

contextOptional[List[str]]Required

This is the context of the test case.

Example: ["Everest is 8,848 metres tall."]

retrieval_contextOptional[List[str]]Required

This is the retrieval context of the test case.

Example: ["Everest is 8,848 metres tall."]

tools_calledOptional[List[ToolCall]]Required

This is the tools called of the test case.

See ToolCall.

expected_toolsOptional[List[ToolCall]]Required

This is the expected tools of the test case.

See ToolCall.

idstrRequired

This is the id of the test case generated by Confident AI.

Example: "<TEST-CASE-ID>"

namestrRequired

This is the name of the test case.

Example: "everest-height"

successOptional[bool]Required

Whether this test case passed all metric thresholds, or null while it is still being evaluated.

Example: true

run_durationOptional[float]Required

The duration of the test case evaluation in seconds.

Example: 1.2

evaluation_costOptional[float]Required

The cost of evaluating this test case.

Example: 0.001

commentsOptional[str]Required

Any comments associated with this test case.

Example: "Reviewed by the geography team."

additional_metadataOptional[Dict[str, Any]]Required

Additional metadata associated with this test case.

Example: {"region":"Nepal"}

metrics_dataList[TestCaseMetricData]Required

The metric evaluation results for this test case.

See TestCaseMetricData.

ToolCall

A tool your LLM application invoked, with what it passed in and what came back.

class ToolCall:
    name: str
    type: Optional[ToolCallType] = None
    description: Optional[str] = None
    input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
    output: Optional[Any] = None
    reasoning: Optional[str] = None

namestrRequired

This is the name of the tool.

Example: "get_landmark_info"

typeOptional[ToolCallType]

descriptionOptional[str]

This is the description of the tool.

Example: "This tool gives information about a mountain."

input_parametersOptional[Dict[str, Any]]

This is the input parameters that are passed to the tool.

Example: {"mountain":"Everest"}

outputOptional[Any]

This is the output of the tool.

Example: "8,848 metres"

reasoningOptional[str]

This is the reasoning your LLM provided for the tool call.

Example: "The user asked for the height of a mountain."

ToolCallType

The type of the tool call, either a function or an MCP tool.

class ToolCallType(Enum):
    FUNCTION = "FUNCTION"
    MCP = "MCP"

FUNCTION · MCP

TopicSummary

The findings for one topic of test cases.

class TopicSummary:
    topic: str
    summary_points: List[SummaryPoint] = Field(alias="summaryPoints")
    test_case_ids: List[str] = Field(alias="testCaseIds")

topicstrRequired

The topic the test cases were grouped under.

Example: "Mountain heights"

summary_pointsList[SummaryPoint]Required

The findings for this topic.

See SummaryPoint.

test_case_idsList[str]Required

The ids of the test cases grouped under this topic.

Example: ["<TEST-CASE-ID>"]

Turn

One message in a conversation, from either the user or the assistant, with the context and tools behind an assistant reply.

class Turn:
    id: Optional[str] = None
    role: TurnRole
    content: str
    user_id: Optional[str] = Field(default=None, alias="userId")
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")

idOptional[str]

The id of a turn assigned by Confident AI.

Example: "<TURN-ID>"

roleTurnRoleRequired

contentstrRequired

The message content of the turn.

Example: "How tall is Mount Everest?"

user_idOptional[str]

The user ID associated with the turn.

Example: "end-user-42"

retrieval_contextOptional[List[str]]

The contexts retrieved to generate the LLM response for this turn.

Example: ["Everest is 8,848 metres tall."]

tools_calledOptional[List[ToolCall]]

The tools called to generate the LLM response for this turn.

See ToolCall.

TurnRole

The role of the turn, either user or assistant.

class TurnRole(Enum):
    USER = "user"
    ASSISTANT = "assistant"

USER · ASSISTANT

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI