Launch Week 3: Five days of launches

Evaluate

Overview

The Confident AI SDK exposes every Evaluate method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

Run Evals

Runs the metrics in metricCollection against your test cases and returns the test run id they were evaluated in. Send either single-turn test cases or multi-turn test cases, not both.

from confident_ai import ConfidentAI
from confident_ai.evaluate import SingleTurnTestCase
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

result = client.evaluate.run_evals(
    metric_collection="Collection Name",
    test_cases=[
        SingleTurnTestCase(
            input="How tall is mount everest?",
            actual_output="No clue, pretty tall I guess?",
            expected_output="Mount Everest is 8,848 metres tall.",
            retrieval_context=["Everest is 8,848 metres tall."],
            tools_called=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            expected_tools=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            context=["Everest is 8,848 metres tall."],
            token_cost=0.002,
            input_token_count=24,
            output_token_count=12,
            name="everest-height",
            flaky=False,
            images_mapping={
                "summit": {"url": "https://example.com/everest.png", "local": False}
            },
            additional_metadata={"region": "Nepal"},
            custom_column_key_values={"team": "search"},
            tags=["geography"]
        )
    ],
    hyperparameters={"model": "gpt-4o-mini"},
    identifier="run-399-102",
)

For async mode, call a_run_evals and await it as shown below:

result = await client.evaluate.a_run_evals(...)

Parameters

ParameterTypeDescription
metric_collectionstrRequired. The name of the metric collection you wish to use for evaluation.
test_casesList[TestCase]Required. This is the list of test cases to evaluate. Every test case in one request must be of the same kind — all single-turn, or all multi-turn. See TestCase.
hyperparametersOptional[Dict[str, HyperparameterValue]]This is any hyperparameters like model or prompt you wish to associate with the test run. See HyperparameterValue.
identifierOptional[str]A unique identifier for the test run.

Returns

This method returns an object of type EvaluateResult.

Evaluate Span

Queues an evaluation of a span against the metrics in metricCollection. The evaluation runs in the background, and its results are stored on the span, so fetch the span to read them once it has finished.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.evaluate.evaluate_span(
    span_uuid="<SPAN-UUID>",
    metric_collection="Collection Name",
    overwrite_metrics=False,
)

For async mode, call a_evaluate_span and await it as shown below:

result = await client.evaluate.a_evaluate_span(...)

Parameters

ParameterTypeDescription
span_uuidstrRequired. The unique identifier of the span.
metric_collectionstrRequired. The name of the single-turn metric collection you wish to use for evaluation.
overwrite_metricsOptional[bool]Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet.

Returns

This method returns an object of type EvaluateSpanResult.

Evaluate Thread

Queues an evaluation of a thread against the multi-turn metrics in metricCollection. The evaluation runs in the background, and its results are stored on the thread, so fetch the thread to read them once it has finished.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.evaluate.evaluate_thread(
    thread_id="thread-42",
    metric_collection="Collection Name",
    chatbot_role="A helpful geography assistant.",
    overwrite_metrics=False,
)

For async mode, call a_evaluate_thread and await it as shown below:

result = await client.evaluate.a_evaluate_thread(...)

Parameters

ParameterTypeDescription
thread_idstrRequired. The id of the thread, as you supplied it when creating its traces.
metric_collectionstrRequired. The name of the multi-turn metric collection you wish to use for evaluation.
chatbot_roleOptional[str]This is the role of the chatbot in the thread, which the multi-turn metrics that judge role adherence evaluate the thread against.
overwrite_metricsOptional[bool]Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet.

Returns

This method returns an object of type EvaluateThreadResult.

Evaluate Trace

Queues an evaluation of a trace against the metrics in metricCollection. The evaluation runs in the background, and its results are stored on the trace, so fetch the trace to read them once it has finished.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.evaluate.evaluate_trace(
    trace_uuid="<TRACE-UUID>",
    metric_collection="Collection Name",
    overwrite_metrics=False,
)

For async mode, call a_evaluate_trace and await it as shown below:

result = await client.evaluate.a_evaluate_trace(...)

Parameters

ParameterTypeDescription
trace_uuidstrRequired. The unique identifier of the trace.
metric_collectionstrRequired. The name of the single-turn metric collection you wish to use for evaluation.
overwrite_metricsOptional[bool]Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet.

Returns

This method returns an object of type EvaluateTraceResult.

Types

EvaluateResult

class EvaluateResult:
    id: str

idstrRequired

This is the unique ID for the test run. This ID is generated by Confident AI and is not to be confused with the identifier provided by the user.

Example: "<TEST-RUN-ID>"

EvaluateSpanResult

The span whose evaluation was queued.

class EvaluateSpanResult:
    id: str

idstrRequired

This is the uuid of the span the evaluation was queued for.

Example: "<SPAN-UUID>"

EvaluateThreadResult

The thread whose evaluation was queued.

class EvaluateThreadResult:
    id: str

idstrRequired

This is the id of the thread the evaluation was queued for.

Example: "thread-42"

EvaluateTraceResult

The trace whose evaluation was queued.

class EvaluateTraceResult:
    id: str

idstrRequired

This is the uuid of the trace the evaluation was queued for.

Example: "<TRACE-UUID>"

HyperparameterValue

A plain value such as a model name, or a reference to the prompt the run used.

HyperparameterValue = Union[
    PromptHyperparameter,
]

A HyperparameterValue is one of the shapes below. Send the fields of one of them, never a mix of both.

A reference to the prompt version the run used.

class PromptHyperparameter:
    id: str
    type: PromptType

idstrRequired

This is the id of the prompt version the run used.

Example: "<PROMPT-VERSION-ID>"

typePromptTypeRequired

MLLMImage

An image referenced from a text field by a [DEEPEVAL:IMAGE:<key>] marker. Send either a public url or the bytes in base64.

class MLLMImage:
    url: str
    local: bool
    base64: Optional[str] = None
    filename: Optional[str] = None
    mime_type: Optional[str] = Field(default=None, alias="mimeType")
    data_base64: Optional[str] = Field(default=None, alias="dataBase64")

urlstrRequired

This is the URL of the image.

Example: "https://example.com/everest.png"

localboolRequired

This is true when the image is your local file.

Example: false

base64Optional[str]

The base64 data of the image.

Example: "iVBORw0KGgo="

filenameOptional[str]

The original file name.

Example: "everest.png"

mime_typeOptional[str]

The image's MIME type.

Example: "image/png"

data_base64Optional[str]

The image encoded as a base64 data URL.

Example: "data:image/png;base64,iVBORw0KGgo="

PromptType

This is the type of the prompt, which can be either a simple text or a list of messages.

class PromptType(Enum):
    TEXT = "TEXT"
    LIST = "LIST"

TEXT · LIST

TestCase

One test case to evaluate: single-turn when it carries input, multi-turn when it carries turns. A test case cannot be both, and one request cannot mix the two kinds.

TestCase = Union[
    SingleTurnTestCase,
    MultiTurnTestCase,
]

A TestCase is one of the shapes below. Send the fields of one of them, never a mix of both.

A test case for a single exchange with your LLM application.

class SingleTurnTestCase:
    input: str
    actual_output: Optional[str] = Field(default=None, alias="actualOutput")
    expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
    context: Optional[List[str]] = None
    token_cost: Optional[float] = Field(default=None, alias="tokenCost")
    input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
    output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
    name: Optional[str] = None
    flaky: Optional[bool] = None
    images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
    additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
    custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
    tags: Optional[List[str]] = None

inputstrRequired

This is the input to your LLM application.

Example: "How tall is mount everest?"

actual_outputOptional[str]

This is the actual output of your LLM application.

Example: "No clue, pretty tall I guess?"

expected_outputOptional[str]

This is the expected output of your LLM application, which is the ideal actual output.

Example: "Mount Everest is 8,848 metres tall."

retrieval_contextOptional[List[str]]

This is the retrieval context of your LLM application.

Example: ["Everest is 8,848 metres tall."]

tools_calledOptional[List[ToolCall]]

This is the tools called by your LLM application.

See ToolCall.

expected_toolsOptional[List[ToolCall]]

This is the expected tools to be called by the LLM application.

See ToolCall.

contextOptional[List[str]]

This is the ideal retrieval context of your LLM application.

Example: ["Everest is 8,848 metres tall."]

token_costOptional[float]

This is the cost of the tokens used by the LLM model.

Example: 0.002

input_token_countOptional[int]

This is the number of input tokens passed to the LLM model.

Example: 24

output_token_countOptional[int]

This is the number of output tokens generated by the LLM model.

Example: 12

nameOptional[str]

This is the name of your test case, it allows you to search and match test cases across different test runs.

Example: "everest-height"

flakyOptional[bool]

This is true if the test case's verdict was non-deterministic across runs.

Example: false

images_mappingOptional[Dict[str, MLLMImage]]

This is the mapping of image placeholders in your test case to the images they refer to.

See MLLMImage.

Example: {"summit":{"url":"https://example.com/everest.png","local":false}}

additional_metadataOptional[Dict[str, Any]]

Additional metadata associated with this test case.

Example: {"region":"Nepal"}

custom_column_key_valuesOptional[Dict[str, str]]

This is the custom column key values of the LLM application.

Example: {"team":"search"}

tagsOptional[List[str]]

This is the list of tags associated with the test case, which is useful for grouping and filtering for test cases.

Example: ["geography"]

ToolCall

A tool your LLM application invoked, with what it passed in and what came back.

class ToolCall:
    name: str
    type: Optional[ToolCallType] = None
    description: Optional[str] = None
    input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
    output: Optional[Any] = None
    reasoning: Optional[str] = None

namestrRequired

This is the name of the tool.

Example: "get_landmark_info"

typeOptional[ToolCallType]

descriptionOptional[str]

This is the description of the tool.

Example: "This tool gives information about a mountain."

input_parametersOptional[Dict[str, Any]]

This is the input parameters that are passed to the tool.

Example: {"mountain":"Everest"}

outputOptional[Any]

This is the output of the tool.

Example: "8,848 metres"

reasoningOptional[str]

This is the reasoning your LLM provided for the tool call.

Example: "The user asked for the height of a mountain."

ToolCallType

The type of the tool call, either a function or an MCP tool.

class ToolCallType(Enum):
    FUNCTION = "FUNCTION"
    MCP = "MCP"

FUNCTION · MCP

Turn

One message in a conversation, from either the user or the assistant, with the context and tools behind an assistant reply.

class Turn:
    id: Optional[str] = None
    role: TurnRole
    content: str
    user_id: Optional[str] = Field(default=None, alias="userId")
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")

idOptional[str]

The id of a turn assigned by Confident AI.

Example: "<TURN-ID>"

roleTurnRoleRequired

contentstrRequired

The message content of the turn.

Example: "How tall is Mount Everest?"

user_idOptional[str]

The user ID associated with the turn.

Example: "end-user-42"

retrieval_contextOptional[List[str]]

The contexts retrieved to generate the LLM response for this turn.

Example: ["Everest is 8,848 metres tall."]

tools_calledOptional[List[ToolCall]]

The tools called to generate the LLM response for this turn.

See ToolCall.

TurnRole

The role of the turn, either user or assistant.

class TurnRole(Enum):
    USER = "user"
    ASSISTANT = "assistant"

USER · ASSISTANT

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI