Launch Week 3: Five days of launches

Datasets

Every Datasets method in the Confident AI Python and TypeScript SDKs.

Overview

The Confident AI SDK exposes every Dataset method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Dataset

client.dataset() returns a Dataset object that stands for one dataset. This object stores the fields listed below, and passes the dataset's id to every method called on it, so you don't need to pass the id nor the stored fields as arguments.

push is the exception, because it sends the stored fields rather than calling a route that names the id. Open the object with alias to call it.

from confident_ai import ConfidentAI

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")

Properties

These are the fields a Dataset stores. A method that loads the dataset fills them in, and a method that saves it sends whichever of them you have set, so set them before you save and read them after you load.

ParameterTypeDescription
dataset_idOptional[str]The unique id of the dataset.
idOptional[str]This is the unique id of the dataset.
aliasOptional[str]This is the alias of the dataset, which is unique within your project.
multi_turnOptional[bool]This is true if the dataset is multi-turn, which contains multi-turn goldens. Single-turn datasets have multiTurn set to false and contain single-turn goldens.
versionOptional[str]The version number of the goldens returned, or null when the dataset has no versions.
goldensOptional[List[Golden]]The goldens in the dataset, oldest first. Every golden is single-turn or multi-turn according to the dataset's multiTurn. See Golden.

Methods

Delete Dataset

Permanently deletes the dataset and everything in it: its goldens, versions, tags and custom columns. This action cannot be undone.

from confident_ai import ConfidentAI

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.delete()

For async mode, call a_delete and await it as shown below:

result = await dataset.a_delete(...)

Returns

This method returns an object of type DatasetRef.

Run Evaluation

Starts an evaluation of the dataset's finalized goldens against a metric collection and returns the test run it is evaluated in. It runs asynchronously, so this returns as soon as the run is created. By default the goldens' stored actual outputs are evaluated; supply aiConnectionId or promptAlias, never both, to generate them first.

from confident_ai import ConfidentAI
from confident_ai.common import GenerationMode

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.run_evaluation(
    metric_collection="Answer Quality",
    identifier="Nightly regression",
    version="00.00.01",
    ai_connection_id="<AI-CONNECTION-ID>",
    prompt_alias="capital-lookup",
    prompt_commit="bab04ce",
    generation_mode=GenerationMode.AI_CONNECTION,
    variables_mapping={"question": "Input"},
    include_simulation=False,
    max_concurrent_generation=5,
    generation_timeout=60,
    num_generations=1,
    mcp_server_ids=["<MCP-SERVER-ID>"],
)

For async mode, call a_run_evaluation and await it as shown below:

result = await dataset.a_run_evaluation(...)

Parameters

ParameterTypeDescription
metric_collectionstrRequired. The name of the metric collection to evaluate against. Names come from the list metric collections endpoint.
identifierOptional[str]A label for the resulting test run, used to recognise it in the test runs list.
versionOptional[str]The dataset version to evaluate. Omit this field to evaluate the latest version.
ai_connection_idOptional[str]The id of the AI connection used to generate the actual outputs before evaluating them. Required when generationMode is AI_CONNECTION, and not allowed together with promptAlias.
prompt_aliasOptional[str]The alias of the prompt used to generate the actual outputs before evaluating them. Required when generationMode is PROMPT, and not allowed together with aiConnectionId.
prompt_commitOptional[str]The prompt commit hash to generate with. Requires promptAlias. Omit this field to generate with the latest commit on the prompt's main branch.
generation_modeOptional[GenerationMode]See GenerationMode.
variables_mappingOptional[Dict[str, str]]Maps each variable in the prompt to the golden field it is interpolated with, such as Input or Expected Output, or to a dataset custom column key. This field applies only when generating from a prompt.
include_simulationOptional[bool]Whether to simulate a conversation for each golden before evaluating it, for multi-turn datasets. Every golden needs a scenario when this is enabled, and turns when it is disabled.
max_concurrent_generationOptional[int]The maximum number of generation calls to run in parallel. An AI connection's own maxConcurrency takes precedence over this value.
generation_timeoutOptional[int]The number of seconds to wait for a single generation before it is marked as errored.
num_generationsOptional[int]How many times to run each golden, so a single outlier response does not skew the results. Omit this field to use the AI connection's defaultNumGenerations, which is 1 when the connection does not set one.
mcp_server_idsOptional[List[str]]The ids of the MCP servers to attach to the run. A tool call whose name matches a tool exposed by one of these servers is labeled an MCP tool call rather than a function call.

Returns

This method returns an object of type RunDatasetEvaluationResult.

Single Turn Golden

Adds goldens to the dataset as unfinalized goldens, for review on the platform before they are used in evaluations.

A single-turn golden to write: one input to your LLM application and the outputs expected of it.

from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.queue_goldens(
    goldens=[
        SingleTurnGoldenRequest(
            input="What is the capital of France?",
            actual_output="The capital of France is Paris.",
            expected_output="Paris.",
            context=["Paris is the capital of France."],
            retrieval_context=[
                "Paris is the capital and largest city of France."
            ],
            tools_called=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            expected_tools=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            token_cost=0.002,
            input_token_count=12,
            output_token_count=3,
            additional_metadata={"source": "faq"},
            comments="Reviewed by the support team.",
            source_file="capitals.csv",
            source_files=["capitals.csv"],
            finalized=True,
            custom_column_key_values={"difficulty": "easy"},
            images_mapping={
                "map": {"url": "https://example.com/paris.png", "local": False}
            },
            tags=["geography"]
        )
    ],
)

For async mode, call a_queue_goldens and await it as shown below:

result = await dataset.a_queue_goldens(...)

Parameters

ParameterTypeDescription
goldensList[GoldenRequest]Required. The goldens to queue for review. Every golden in one request must be of the same kind and match the dataset's multiTurn. They are stored unfinalized, whatever each golden's own finalized says. See GoldenRequest.

Returns

This method returns an object of type DatasetRef.

Multi Turn Golden

Adds goldens to the dataset as unfinalized goldens, for review on the platform before they are used in evaluations.

A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.

from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.queue_goldens(
    goldens=[
        MultiTurnGoldenRequest(
            scenario="A traveller wants to book a hotel in Paris.",
            expected_outcome="The assistant confirms a reservation near the Louvre.",
            user_description="A traveller planning a weekend in Paris.",
            turns=[
                {
                    "role": "user",
                    "content": "I need a hotel in Paris near the Louvre."
                },
                {
                    "role": "assistant",
                    "content": "Hôtel du Louvre has rooms available. Which dates?"
                }
            ],
            context=[
                "Hôtel du Louvre is a five-minute walk from the museum."
            ],
            additional_metadata={"source": "faq"},
            comments="Reviewed by the support team.",
            source_file="capitals.csv",
            source_files=["capitals.csv"],
            finalized=True,
            custom_column_key_values={"difficulty": "easy"},
            images_mapping={
                "map": {"url": "https://example.com/paris.png", "local": False}
            },
            tags=["geography"]
        )
    ],
)

For async mode, call a_queue_goldens and await it as shown below:

result = await dataset.a_queue_goldens(...)

Parameters

ParameterTypeDescription
goldensList[GoldenRequest]Required. The goldens to queue for review. Every golden in one request must be of the same kind and match the dataset's multiTurn. They are stored unfinalized, whatever each golden's own finalized says. See GoldenRequest.

Returns

This method returns an object of type DatasetRef.

Pull Dataset

Retrieves the dataset with its goldens, oldest first. Pass version to pull a specific version, and finalized=false to pull the goldens still awaiting review instead of the finalized ones. Requires an active trial or paid plan, and the Team plan or above to pull a version.

from confident_ai import ConfidentAI

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.pull(version="00.00.01", finalized="true")

For async mode, call a_pull and await it as shown below:

result = await dataset.a_pull(...)

Parameters

ParameterTypeDescription
versionOptional[str]The version to pull. Defaults to the latest version, or to the unversioned goldens when the dataset has no versions. Requires the Team plan or above.
finalizedOptional[Literal['true', 'false']]Whether to pull the finalized goldens, or with false the goldens still awaiting review. Defaults to true.

Returns

This method returns an object of type Dataset.

Push Dataset

Adds goldens to a dataset, creating it when the alias names none yet, and returns the dataset's id. Pushing to a version requires the Team plan or above.

from confident_ai import ConfidentAI

client = ConfidentAI()

dataset = client.dataset(alias="capitals")
result = dataset.push(finalized=True)

For async mode, call a_push and await it as shown below:

result = await dataset.a_push(...)

Parameters

ParameterTypeDescription
finalizedOptional[bool]Whether the goldens pushed are finalized, that is ready to use in evaluations. Applies to every golden in this request.

Returns

This method returns an object of type DatasetRef.

Methods (Stateless)

These methods take every argument themselves, so a caller reaches them through client.datasets without opening a Dataset first.

List Datasets

Lists all the datasets in your Confident AI project, newest first, without their goldens.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.list()

For async mode, call a_list and await it as shown below:

result = await client.datasets.a_list(...)

Returns

This method returns an object of type DatasetList.

Push Dataset

Adds goldens to a dataset, creating it when the alias names none yet, and returns the dataset's id. Pushing to a version requires the Team plan or above.

from confident_ai import ConfidentAI
from confident_ai.datasets import PushSingleTurnGolden
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

result = client.datasets.push(
    alias="capitals",
    goldens=[
        PushSingleTurnGolden(
            input="What is the capital of France?",
            actual_output="The capital of France is Paris.",
            expected_output="Paris.",
            context=["Paris is the capital of France."],
            retrieval_context=[
                "Paris is the capital and largest city of France."
            ],
            tools_called=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            expected_tools=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            token_cost=0.002,
            input_token_count=12,
            output_token_count=3,
            additional_metadata={"source": "faq"},
            comments="Reviewed by the support team.",
            source_file="capitals.csv",
            source_files=["capitals.csv"],
            finalized=True,
            custom_column_key_values={"difficulty": "easy"},
            images_mapping={
                "map": {"url": "https://example.com/paris.png", "local": False}
            },
            tags=["geography"],
            id="<GOLDEN-ID>"
        )
    ],
    finalized=True,
    version="00.00.01",
)

For async mode, call a_push and await it as shown below:

result = await client.datasets.a_push(...)

Parameters

ParameterTypeDescription
aliasstrRequired. The alias of the dataset, unique within your project. A new dataset is created when no dataset with this alias exists.
goldensList[PushGolden]Required. The goldens to push. A golden carrying an id updates the golden already in the dataset; one without is added. Every golden in one request must be of the same kind — all single- turn, or all multi-turn — and match the dataset's multiTurn. A new dataset takes its kind from them. See PushGolden.
finalizedOptional[bool]Whether the goldens pushed are finalized, that is ready to use in evaluations. Applies to every golden in this request.
versionOptional[str]The dataset version to push the goldens onto, for example 00.00.01. When the dataset has versions, omitting it pushes to the latest version; when it has none, omitting it leaves the goldens unversioned. A version cannot be given for a dataset that does not exist yet. Requires the Team plan or above.

Returns

This method returns an object of type DatasetRef.

Pull Dataset

Retrieves the dataset with its goldens, oldest first. Pass version to pull a specific version, and finalized=false to pull the goldens still awaiting review instead of the finalized ones. Requires an active trial or paid plan, and the Team plan or above to pull a version.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.pull(
    dataset_id="<DATASET-ID>",
    version="00.00.01",
    finalized="true",
)

For async mode, call a_pull and await it as shown below:

result = await client.datasets.a_pull(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
versionOptional[str]The version to pull. Defaults to the latest version, or to the unversioned goldens when the dataset has no versions. Requires the Team plan or above.
finalizedOptional[Literal['true', 'false']]Whether to pull the finalized goldens, or with false the goldens still awaiting review. Defaults to true.

Returns

This method returns an object of type Dataset.

Delete Dataset

Permanently deletes the dataset and everything in it: its goldens, versions, tags and custom columns. This action cannot be undone.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.delete(dataset_id="<DATASET-ID>")

For async mode, call a_delete and await it as shown below:

result = await client.datasets.a_delete(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.

Returns

This method returns an object of type DatasetRef.

Single Turn Golden

Adds goldens to the dataset as unfinalized goldens, for review on the platform before they are used in evaluations.

A single-turn golden to write: one input to your LLM application and the outputs expected of it.

from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

result = client.datasets.queue_goldens(
    dataset_id="<DATASET-ID>",
    goldens=[
        SingleTurnGoldenRequest(
            input="What is the capital of France?",
            actual_output="The capital of France is Paris.",
            expected_output="Paris.",
            context=["Paris is the capital of France."],
            retrieval_context=[
                "Paris is the capital and largest city of France."
            ],
            tools_called=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            expected_tools=[
                ToolCall(
                    name="get_landmark_info",
                    type=ToolCallType.FUNCTION,
                    description="This tool gives information about a mountain.",
                    input_parameters={"mountain": "Everest"},
                    output="8,848 metres",
                    reasoning="The user asked for the height of a mountain."
                )
            ],
            token_cost=0.002,
            input_token_count=12,
            output_token_count=3,
            additional_metadata={"source": "faq"},
            comments="Reviewed by the support team.",
            source_file="capitals.csv",
            source_files=["capitals.csv"],
            finalized=True,
            custom_column_key_values={"difficulty": "easy"},
            images_mapping={
                "map": {"url": "https://example.com/paris.png", "local": False}
            },
            tags=["geography"]
        )
    ],
)

For async mode, call a_queue_goldens and await it as shown below:

result = await client.datasets.a_queue_goldens(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
goldensList[GoldenRequest]Required. The goldens to queue for review. Every golden in one request must be of the same kind and match the dataset's multiTurn. They are stored unfinalized, whatever each golden's own finalized says. See GoldenRequest.

Returns

This method returns an object of type DatasetRef.

Multi Turn Golden

Adds goldens to the dataset as unfinalized goldens, for review on the platform before they are used in evaluations.

A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.

from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest

client = ConfidentAI()

result = client.datasets.queue_goldens(
    dataset_id="<DATASET-ID>",
    goldens=[
        MultiTurnGoldenRequest(
            scenario="A traveller wants to book a hotel in Paris.",
            expected_outcome="The assistant confirms a reservation near the Louvre.",
            user_description="A traveller planning a weekend in Paris.",
            turns=[
                {
                    "role": "user",
                    "content": "I need a hotel in Paris near the Louvre."
                },
                {
                    "role": "assistant",
                    "content": "Hôtel du Louvre has rooms available. Which dates?"
                }
            ],
            context=[
                "Hôtel du Louvre is a five-minute walk from the museum."
            ],
            additional_metadata={"source": "faq"},
            comments="Reviewed by the support team.",
            source_file="capitals.csv",
            source_files=["capitals.csv"],
            finalized=True,
            custom_column_key_values={"difficulty": "easy"},
            images_mapping={
                "map": {"url": "https://example.com/paris.png", "local": False}
            },
            tags=["geography"]
        )
    ],
)

For async mode, call a_queue_goldens and await it as shown below:

result = await client.datasets.a_queue_goldens(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
goldensList[GoldenRequest]Required. The goldens to queue for review. Every golden in one request must be of the same kind and match the dataset's multiTurn. They are stored unfinalized, whatever each golden's own finalized says. See GoldenRequest.

Returns

This method returns an object of type DatasetRef.

Run Evaluation

Starts an evaluation of the dataset's finalized goldens against a metric collection and returns the test run it is evaluated in. It runs asynchronously, so this returns as soon as the run is created. By default the goldens' stored actual outputs are evaluated; supply aiConnectionId or promptAlias, never both, to generate them first.

from confident_ai import ConfidentAI
from confident_ai.common import GenerationMode

client = ConfidentAI()

result = client.datasets.run_evaluation(
    dataset_id="<DATASET-ID>",
    metric_collection="Answer Quality",
    identifier="Nightly regression",
    version="00.00.01",
    ai_connection_id="<AI-CONNECTION-ID>",
    prompt_alias="capital-lookup",
    prompt_commit="bab04ce",
    generation_mode=GenerationMode.AI_CONNECTION,
    variables_mapping={"question": "Input"},
    include_simulation=False,
    max_concurrent_generation=5,
    generation_timeout=60,
    num_generations=1,
    mcp_server_ids=["<MCP-SERVER-ID>"],
)

For async mode, call a_run_evaluation and await it as shown below:

result = await client.datasets.a_run_evaluation(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
metric_collectionstrRequired. The name of the metric collection to evaluate against. Names come from the list metric collections endpoint.
identifierOptional[str]A label for the resulting test run, used to recognise it in the test runs list.
versionOptional[str]The dataset version to evaluate. Omit this field to evaluate the latest version.
ai_connection_idOptional[str]The id of the AI connection used to generate the actual outputs before evaluating them. Required when generationMode is AI_CONNECTION, and not allowed together with promptAlias.
prompt_aliasOptional[str]The alias of the prompt used to generate the actual outputs before evaluating them. Required when generationMode is PROMPT, and not allowed together with aiConnectionId.
prompt_commitOptional[str]The prompt commit hash to generate with. Requires promptAlias. Omit this field to generate with the latest commit on the prompt's main branch.
generation_modeOptional[GenerationMode]See GenerationMode.
variables_mappingOptional[Dict[str, str]]Maps each variable in the prompt to the golden field it is interpolated with, such as Input or Expected Output, or to a dataset custom column key. This field applies only when generating from a prompt.
include_simulationOptional[bool]Whether to simulate a conversation for each golden before evaluating it, for multi-turn datasets. Every golden needs a scenario when this is enabled, and turns when it is disabled.
max_concurrent_generationOptional[int]The maximum number of generation calls to run in parallel. An AI connection's own maxConcurrency takes precedence over this value.
generation_timeoutOptional[int]The number of seconds to wait for a single generation before it is marked as errored.
num_generationsOptional[int]How many times to run each golden, so a single outlier response does not skew the results. Omit this field to use the AI connection's defaultNumGenerations, which is 1 when the connection does not set one.
mcp_server_idsOptional[List[str]]The ids of the MCP servers to attach to the run. A tool call whose name matches a tool exposed by one of these servers is labeled an MCP tool call rather than a function call.

Returns

This method returns an object of type RunDatasetEvaluationResult.

Types

Dataset

A pulled dataset with the goldens of one version.

class Dataset:
    id: str
    alias: str
    multi_turn: bool = Field(alias="multiTurn")
    version: Optional[str]
    goldens: List[Golden]

idstrRequired

This is the unique id of the dataset.

Example: "<DATASET-ID>"

aliasstrRequired

This is the alias of the dataset, which is unique within your project.

Example: "capitals"

multi_turnboolRequired

This is true if the dataset is multi-turn, which contains multi-turn goldens. Single-turn datasets have multiTurn set to false and contain single-turn goldens.

Example: false

versionOptional[str]Required

The version number of the goldens returned, or null when the dataset has no versions.

Example: "00.00.01"

goldensList[Golden]Required

The goldens in the dataset, oldest first. Every golden is single-turn or multi-turn according to the dataset's multiTurn.

See Golden.

DatasetList

class DatasetList:
    datasets: List[DatasetSummary]

datasetsList[DatasetSummary]Required

This is the list of datasets in your project, newest first.

See DatasetSummary.

DatasetRef

class DatasetRef:
    id: str

idstrRequired

This is the unique id of the dataset.

Example: "<DATASET-ID>"

DatasetSummary

A dataset as it appears in your project's dataset list, without any goldens.

class DatasetSummary:
    id: str
    alias: str
    multi_turn: bool = Field(alias="multiTurn")

idstrRequired

This is the unique id of the dataset.

Example: "<DATASET-ID>"

aliasstrRequired

This is the alias of the dataset, which is unique within your project.

Example: "capitals"

multi_turnboolRequired

This is true if the dataset is multi-turn, which contains multi-turn goldens. Single-turn datasets have multiTurn set to false and contain single-turn goldens.

Example: false

GenerationMode

Where the actual outputs come from when running a dataset: AI_CONNECTION generates them with an AI connection, PROMPT with a prompt. Omit it when you supply at most one of aiConnectionId or promptAlias, and Confident AI infers the mode from whichever you sent.

class GenerationMode(Enum):
    AI_CONNECTION = "AI_CONNECTION"
    PROMPT = "PROMPT"

AI_CONNECTION · PROMPT

Golden

A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is.

Golden = Union[
    SingleTurnGolden,
    MultiTurnGolden,
]

A Golden is one of the shapes below. Send the fields of one of them, never a mix of both.

A single-turn golden as stored in the dataset.

class SingleTurnGolden:
    input: str
    actual_output: Optional[str] = Field(default=None, alias="actualOutput")
    expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
    context: Optional[List[str]] = None
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
    token_cost: Optional[float] = Field(default=None, alias="tokenCost")
    input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
    output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
    id: Optional[str] = None
    additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
    comments: Optional[str] = None
    source_file: Optional[str] = Field(default=None, alias="sourceFile")
    source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
    finalized: Optional[bool] = None
    custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
    tags: Optional[List[str]] = None

inputstrRequired

This is the input to your LLM application.

Example: "What is the capital of France?"

actual_outputOptional[str]

This is the actual output of your LLM application.

Example: "The capital of France is Paris."

expected_outputOptional[str]

This is the expected output of your LLM application, which is the ideal actual output.

Example: "Paris."

contextOptional[List[str]]

This is the ideal retrieval context of your LLM application.

Example: ["Paris is the capital of France."]

retrieval_contextOptional[List[str]]

This is the retrieval context of your LLM application.

Example: ["Paris is the capital and largest city of France."]

tools_calledOptional[List[ToolCall]]

This is the tools called by your LLM application.

See ToolCall.

expected_toolsOptional[List[ToolCall]]

This is the expected tools to be called by the LLM application.

See ToolCall.

token_costOptional[float]

This is the cost of the tokens used to produce the actual output.

Example: 0.002

input_token_countOptional[int]

This is the number of input tokens passed to the LLM model.

Example: 12

output_token_countOptional[int]

This is the number of output tokens generated by the LLM model.

Example: 3

idOptional[str]

The id of the golden assigned by Confident AI. Use it to get, update or delete this golden.

Example: "<GOLDEN-ID>"

additional_metadataOptional[Dict[str, Any]]

This is any additional metadata associated with the golden.

Example: {"source":"faq"}

commentsOptional[str]

This is any comments associated with the golden.

Example: "Reviewed by the support team."

source_fileOptional[str]

This is the source file from which the golden was retrieved.

Example: "capitals.csv"

source_filesOptional[List[str]]

These are the source files the golden was retrieved from.

Example: ["capitals.csv"]

finalizedOptional[bool]

This is true when the golden is finalized and ready to use in evaluations.

Example: true

custom_column_key_valuesOptional[Dict[str, str]]

Key-value pairs representing custom table column data for this golden. Keys correspond to the custom column keys defined in the dataset. Absent when the golden has no custom column values.

Example: {"difficulty":"easy"}

tagsOptional[List[str]]

These are the tags associated with the golden.

Example: ["geography"]

GoldenRequest

One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn.

GoldenRequest = Union[
    SingleTurnGoldenRequest,
    MultiTurnGoldenRequest,
]

A GoldenRequest is one of the shapes below. Send the fields of one of them, never a mix of both.

A single-turn golden to write: one input to your LLM application and the outputs expected of it.

class SingleTurnGoldenRequest:
    input: str
    actual_output: Optional[str] = Field(default=None, alias="actualOutput")
    expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
    context: Optional[List[str]] = None
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
    token_cost: Optional[float] = Field(default=None, alias="tokenCost")
    input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
    output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
    additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
    comments: Optional[str] = None
    source_file: Optional[str] = Field(default=None, alias="sourceFile")
    source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
    finalized: Optional[bool] = None
    custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
    images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
    tags: Optional[List[str]] = None

inputstrRequired

This is the input to your LLM application.

Example: "What is the capital of France?"

actual_outputOptional[str]

This is the actual output of your LLM application.

Example: "The capital of France is Paris."

expected_outputOptional[str]

This is the expected output of your LLM application, which is the ideal actual output.

Example: "Paris."

contextOptional[List[str]]

This is the ideal retrieval context of your LLM application.

Example: ["Paris is the capital of France."]

retrieval_contextOptional[List[str]]

This is the retrieval context of your LLM application.

Example: ["Paris is the capital and largest city of France."]

tools_calledOptional[List[ToolCall]]

This is the tools called by your LLM application.

See ToolCall.

expected_toolsOptional[List[ToolCall]]

This is the expected tools to be called by the LLM application.

See ToolCall.

token_costOptional[float]

This is the cost of the tokens used to produce the actual output.

Example: 0.002

input_token_countOptional[int]

This is the number of input tokens passed to the LLM model.

Example: 12

output_token_countOptional[int]

This is the number of output tokens generated by the LLM model.

Example: 3

additional_metadataOptional[Dict[str, Any]]

Additional metadata to associate with the golden.

Example: {"source":"faq"}

commentsOptional[str]

Comments to associate with the golden.

Example: "Reviewed by the support team."

source_fileOptional[str]

The source file the golden was retrieved from. Like tags and customColumnKeyValues, this is left unchanged when the request omits it; send null to clear it.

Example: "capitals.csv"

source_filesOptional[List[str]]

The source files the golden was retrieved from. Like tags and customColumnKeyValues, these are left unchanged when the request omits them; send an empty array to clear them.

Example: ["capitals.csv"]

finalizedOptional[bool]

Whether the golden is ready to use in evaluations. When pushing or queueing a list of goldens the request decides this for every golden and this field is ignored.

Example: true

custom_column_key_valuesOptional[Dict[str, str]]

Custom dataset column values keyed by column name. A column that does not exist in the dataset yet is created.

Example: {"difficulty":"easy"}

images_mappingOptional[Dict[str, MLLMImage]]

The media this golden refers to, keyed by the id inside each placeholder. Put [DEEPEVAL:IMAGE:<id>] or [DEEPEVAL:PDF:<id>] in a text field where the media belongs, and the platform substitutes the entry with a matching key.

See MLLMImage.

Example: {"map":{"url":"https://example.com/paris.png","local":false}}

tagsOptional[List[str]]

Tags to associate with the golden, which is useful for grouping and filtering goldens. A tag that does not exist in the dataset yet is created.

Example: ["geography"]

MLLMImage

An image referenced from a text field by a [DEEPEVAL:IMAGE:<key>] marker. Send either a public url or the bytes in base64.

class MLLMImage:
    url: str
    local: bool
    base64: Optional[str] = None
    filename: Optional[str] = None
    mime_type: Optional[str] = Field(default=None, alias="mimeType")
    data_base64: Optional[str] = Field(default=None, alias="dataBase64")

urlstrRequired

This is the URL of the image.

Example: "https://example.com/everest.png"

localboolRequired

This is true when the image is your local file.

Example: false

base64Optional[str]

The base64 data of the image.

Example: "iVBORw0KGgo="

filenameOptional[str]

The original file name.

Example: "everest.png"

mime_typeOptional[str]

The image's MIME type.

Example: "image/png"

data_base64Optional[str]

The image encoded as a base64 data URL.

Example: "data:image/png;base64,iVBORw0KGgo="

PushGolden

One golden to push: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn. Carrying an id updates that golden instead of adding one.

PushGolden = Union[
    PushSingleTurnGolden,
    PushMultiTurnGolden,
]

A PushGolden is one of the shapes below. Send the fields of one of them, never a mix of both.

A single-turn golden to push: one input to your LLM application and the outputs expected of it, updating the golden named by id when there is one.

class PushSingleTurnGolden:
    input: str
    actual_output: Optional[str] = Field(default=None, alias="actualOutput")
    expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
    context: Optional[List[str]] = None
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
    token_cost: Optional[float] = Field(default=None, alias="tokenCost")
    input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
    output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
    additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
    comments: Optional[str] = None
    source_file: Optional[str] = Field(default=None, alias="sourceFile")
    source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
    finalized: Optional[bool] = None
    custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
    images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
    tags: Optional[List[str]] = None
    id: Optional[str] = None

inputstrRequired

This is the input to your LLM application.

Example: "What is the capital of France?"

actual_outputOptional[str]

This is the actual output of your LLM application.

Example: "The capital of France is Paris."

expected_outputOptional[str]

This is the expected output of your LLM application, which is the ideal actual output.

Example: "Paris."

contextOptional[List[str]]

This is the ideal retrieval context of your LLM application.

Example: ["Paris is the capital of France."]

retrieval_contextOptional[List[str]]

This is the retrieval context of your LLM application.

Example: ["Paris is the capital and largest city of France."]

tools_calledOptional[List[ToolCall]]

This is the tools called by your LLM application.

See ToolCall.

expected_toolsOptional[List[ToolCall]]

This is the expected tools to be called by the LLM application.

See ToolCall.

token_costOptional[float]

This is the cost of the tokens used to produce the actual output.

Example: 0.002

input_token_countOptional[int]

This is the number of input tokens passed to the LLM model.

Example: 12

output_token_countOptional[int]

This is the number of output tokens generated by the LLM model.

Example: 3

additional_metadataOptional[Dict[str, Any]]

Additional metadata to associate with the golden.

Example: {"source":"faq"}

commentsOptional[str]

Comments to associate with the golden.

Example: "Reviewed by the support team."

source_fileOptional[str]

The source file the golden was retrieved from. Like tags and customColumnKeyValues, this is left unchanged when the request omits it; send null to clear it.

Example: "capitals.csv"

source_filesOptional[List[str]]

The source files the golden was retrieved from. Like tags and customColumnKeyValues, these are left unchanged when the request omits them; send an empty array to clear them.

Example: ["capitals.csv"]

finalizedOptional[bool]

Whether the golden is ready to use in evaluations. When pushing or queueing a list of goldens the request decides this for every golden and this field is ignored.

Example: true

custom_column_key_valuesOptional[Dict[str, str]]

Custom dataset column values keyed by column name. A column that does not exist in the dataset yet is created.

Example: {"difficulty":"easy"}

images_mappingOptional[Dict[str, MLLMImage]]

The media this golden refers to, keyed by the id inside each placeholder. Put [DEEPEVAL:IMAGE:<id>] or [DEEPEVAL:PDF:<id>] in a text field where the media belongs, and the platform substitutes the entry with a matching key.

See MLLMImage.

Example: {"map":{"url":"https://example.com/paris.png","local":false}}

tagsOptional[List[str]]

Tags to associate with the golden, which is useful for grouping and filtering goldens. A tag that does not exist in the dataset yet is created.

Example: ["geography"]

idOptional[str]

The id of a golden already in this dataset, as returned when the dataset is pulled. That golden is updated in place, staying in the version it is already in. Omit it to add a new golden.

Example: "<GOLDEN-ID>"

RunDatasetEvaluationResult

class RunDatasetEvaluationResult:
    id: str
    test_case_count: int = Field(alias="testCaseCount")

idstrRequired

This is the unique id of the test run the dataset is evaluated in, generated by Confident AI and not to be confused with the identifier you supplied.

Example: "<TEST-RUN-ID>"

test_case_countintRequired

The number of test cases the evaluation was started with.

Example: 42

ToolCall

A tool your LLM application invoked, with what it passed in and what came back.

class ToolCall:
    name: str
    type: Optional[ToolCallType] = None
    description: Optional[str] = None
    input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
    output: Optional[Any] = None
    reasoning: Optional[str] = None

namestrRequired

This is the name of the tool.

Example: "get_landmark_info"

typeOptional[ToolCallType]

descriptionOptional[str]

This is the description of the tool.

Example: "This tool gives information about a mountain."

input_parametersOptional[Dict[str, Any]]

This is the input parameters that are passed to the tool.

Example: {"mountain":"Everest"}

outputOptional[Any]

This is the output of the tool.

Example: "8,848 metres"

reasoningOptional[str]

This is the reasoning your LLM provided for the tool call.

Example: "The user asked for the height of a mountain."

ToolCallType

The type of the tool call, either a function or an MCP tool.

class ToolCallType(Enum):
    FUNCTION = "FUNCTION"
    MCP = "MCP"

FUNCTION · MCP

Turn

One message in a conversation, from either the user or the assistant, with the context and tools behind an assistant reply.

class Turn:
    id: Optional[str] = None
    role: TurnRole
    content: str
    user_id: Optional[str] = Field(default=None, alias="userId")
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")

idOptional[str]

The id of a turn assigned by Confident AI.

Example: "<TURN-ID>"

roleTurnRoleRequired

contentstrRequired

The message content of the turn.

Example: "How tall is Mount Everest?"

user_idOptional[str]

The user ID associated with the turn.

Example: "end-user-42"

retrieval_contextOptional[List[str]]

The contexts retrieved to generate the LLM response for this turn.

Example: ["Everest is 8,848 metres tall."]

tools_calledOptional[List[ToolCall]]

The tools called to generate the LLM response for this turn.

See ToolCall.

TurnRole

The role of the turn, either user or assistant.

class TurnRole(Enum):
    USER = "user"
    ASSISTANT = "assistant"

USER · ASSISTANT

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI