Launch Week 3: Five days of launches

Goldens

Overview

The Confident AI SDK exposes every Golden method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

Single Turn Golden

Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.

A single-turn golden to write: one input to your LLM application and the outputs expected of it.

from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.create_golden(
    golden=SingleTurnGoldenRequest(
        input="What is the capital of France?",
        actual_output="The capital of France is Paris.",
        expected_output="Paris.",
        context=["Paris is the capital of France."],
        retrieval_context=[
            "Paris is the capital and largest city of France."
        ],
        tools_called=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        expected_tools=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        token_cost=0.002,
        input_token_count=12,
        output_token_count=3,
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        images_mapping={
            "map": {"url": "https://example.com/paris.png", "local": False}
        },
        tags=["geography"]
    ),
    version="00.00.01",
)

For async mode, call a_create_golden and await it as shown below:

result = await dataset.a_create_golden(...)

Parameters

ParameterTypeDescription
goldenGoldenRequestRequired. See GoldenRequest.
versionOptional[str]The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions.

Returns

This method returns an object of type GoldenRef.

Multi Turn Golden

Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.

A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.

from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.create_golden(
    golden=MultiTurnGoldenRequest(
        scenario="A traveller wants to book a hotel in Paris.",
        expected_outcome="The assistant confirms a reservation near the Louvre.",
        user_description="A traveller planning a weekend in Paris.",
        turns=[
            {
                "role": "user",
                "content": "I need a hotel in Paris near the Louvre."
            },
            {
                "role": "assistant",
                "content": "Hôtel du Louvre has rooms available. Which dates?"
            }
        ],
        context=["Hôtel du Louvre is a five-minute walk from the museum."],
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        images_mapping={
            "map": {"url": "https://example.com/paris.png", "local": False}
        },
        tags=["geography"]
    ),
    version="00.00.01",
)

For async mode, call a_create_golden and await it as shown below:

result = await dataset.a_create_golden(...)

Parameters

ParameterTypeDescription
goldenGoldenRequestRequired. See GoldenRequest.
versionOptional[str]The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions.

Returns

This method returns an object of type GoldenRef.

Get Golden

Retrieves a single golden by id. It is single-turn or multi-turn according to the dataset's multiTurn.

from confident_ai import ConfidentAI

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.get_golden(golden_id="<GOLDEN-ID>")

For async mode, call a_get_golden and await it as shown below:

result = await dataset.a_get_golden(...)

Parameters

ParameterTypeDescription
golden_idstrRequired. The unique id of the golden, returned when the dataset is pulled.

Returns

This method returns an object of type Golden.

Delete Golden

Permanently deletes a single golden. The rest of the dataset is unchanged, and this action cannot be undone.

from confident_ai import ConfidentAI

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.delete_golden(golden_id="<GOLDEN-ID>")

For async mode, call a_delete_golden and await it as shown below:

result = await dataset.a_delete_golden(...)

Parameters

ParameterTypeDescription
golden_idstrRequired. The unique id of the golden, returned when the dataset is pulled.

Returns

This method returns an object of type GoldenRef.

Single Turn Golden

Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.

A single-turn golden as stored in the dataset.

from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGolden
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.update_golden(
    golden=SingleTurnGolden(
        input="What is the capital of France?",
        actual_output="The capital of France is Paris.",
        expected_output="Paris.",
        context=["Paris is the capital of France."],
        retrieval_context=[
            "Paris is the capital and largest city of France."
        ],
        tools_called=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        expected_tools=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        token_cost=0.002,
        input_token_count=12,
        output_token_count=3,
        id="<GOLDEN-ID>",
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        tags=["geography"]
    ),
)

For async mode, call a_update_golden and await it as shown below:

result = await dataset.a_update_golden(...)

Parameters

ParameterTypeDescription
goldenGoldenRequired. A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is. See GoldenRequest.

Returns

This method returns an object of type GoldenRef.

Multi Turn Golden

Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.

A multi-turn golden as stored in the dataset.

from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGolden

client = ConfidentAI()

dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.update_golden(
    golden=MultiTurnGolden(
        scenario="A traveller wants to book a hotel in Paris.",
        expected_outcome="The assistant confirms a reservation near the Louvre.",
        user_description="A traveller planning a weekend in Paris.",
        turns=[
            {
                "role": "user",
                "content": "I need a hotel in Paris near the Louvre."
            },
            {
                "role": "assistant",
                "content": "Hôtel du Louvre has rooms available. Which dates?"
            }
        ],
        context=["Hôtel du Louvre is a five-minute walk from the museum."],
        id="<GOLDEN-ID>",
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        tags=["geography"]
    ),
)

For async mode, call a_update_golden and await it as shown below:

result = await dataset.a_update_golden(...)

Parameters

ParameterTypeDescription
goldenGoldenRequired. A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is. See GoldenRequest.

Returns

This method returns an object of type GoldenRef.

Methods (Stateless)

These methods take every argument themselves, so a caller reaches them through client.datasets without opening a Dataset first.

Single Turn Golden

Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.

A single-turn golden to write: one input to your LLM application and the outputs expected of it.

from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

result = client.datasets.create_golden(
    dataset_id="<DATASET-ID>",
    golden=SingleTurnGoldenRequest(
        input="What is the capital of France?",
        actual_output="The capital of France is Paris.",
        expected_output="Paris.",
        context=["Paris is the capital of France."],
        retrieval_context=[
            "Paris is the capital and largest city of France."
        ],
        tools_called=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        expected_tools=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        token_cost=0.002,
        input_token_count=12,
        output_token_count=3,
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        images_mapping={
            "map": {"url": "https://example.com/paris.png", "local": False}
        },
        tags=["geography"]
    ),
    version="00.00.01",
)

For async mode, call a_create_golden and await it as shown below:

result = await client.datasets.a_create_golden(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
goldenGoldenRequestRequired. See GoldenRequest.
versionOptional[str]The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions.

Returns

This method returns an object of type GoldenRef.

Multi Turn Golden

Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.

A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.

from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest

client = ConfidentAI()

result = client.datasets.create_golden(
    dataset_id="<DATASET-ID>",
    golden=MultiTurnGoldenRequest(
        scenario="A traveller wants to book a hotel in Paris.",
        expected_outcome="The assistant confirms a reservation near the Louvre.",
        user_description="A traveller planning a weekend in Paris.",
        turns=[
            {
                "role": "user",
                "content": "I need a hotel in Paris near the Louvre."
            },
            {
                "role": "assistant",
                "content": "Hôtel du Louvre has rooms available. Which dates?"
            }
        ],
        context=["Hôtel du Louvre is a five-minute walk from the museum."],
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        images_mapping={
            "map": {"url": "https://example.com/paris.png", "local": False}
        },
        tags=["geography"]
    ),
    version="00.00.01",
)

For async mode, call a_create_golden and await it as shown below:

result = await client.datasets.a_create_golden(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
goldenGoldenRequestRequired. See GoldenRequest.
versionOptional[str]The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions.

Returns

This method returns an object of type GoldenRef.

Get Golden

Retrieves a single golden by id. It is single-turn or multi-turn according to the dataset's multiTurn.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.get_golden(
    dataset_id="<DATASET-ID>",
    golden_id="<GOLDEN-ID>",
)

For async mode, call a_get_golden and await it as shown below:

result = await client.datasets.a_get_golden(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
golden_idstrRequired. The unique id of the golden, returned when the dataset is pulled.

Returns

This method returns an object of type Golden.

Single Turn Golden

Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.

A single-turn golden to write: one input to your LLM application and the outputs expected of it.

from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType

client = ConfidentAI()

result = client.datasets.update_golden(
    dataset_id="<DATASET-ID>",
    golden_id="<GOLDEN-ID>",
    golden=SingleTurnGoldenRequest(
        input="What is the capital of France?",
        actual_output="The capital of France is Paris.",
        expected_output="Paris.",
        context=["Paris is the capital of France."],
        retrieval_context=[
            "Paris is the capital and largest city of France."
        ],
        tools_called=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        expected_tools=[
            ToolCall(
                name="get_landmark_info",
                type=ToolCallType.FUNCTION,
                description="This tool gives information about a mountain.",
                input_parameters={"mountain": "Everest"},
                output="8,848 metres",
                reasoning="The user asked for the height of a mountain."
            )
        ],
        token_cost=0.002,
        input_token_count=12,
        output_token_count=3,
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        images_mapping={
            "map": {"url": "https://example.com/paris.png", "local": False}
        },
        tags=["geography"]
    ),
)

For async mode, call a_update_golden and await it as shown below:

result = await client.datasets.a_update_golden(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
golden_idstrRequired. The unique id of the golden, returned when the dataset is pulled.
goldenGoldenRequestRequired. One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn. Pass a SingleTurnGoldenRequest or a MultiTurnGoldenRequest. See GoldenRequest.

Returns

This method returns an object of type GoldenRef.

Multi Turn Golden

Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.

A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.

from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest

client = ConfidentAI()

result = client.datasets.update_golden(
    dataset_id="<DATASET-ID>",
    golden_id="<GOLDEN-ID>",
    golden=MultiTurnGoldenRequest(
        scenario="A traveller wants to book a hotel in Paris.",
        expected_outcome="The assistant confirms a reservation near the Louvre.",
        user_description="A traveller planning a weekend in Paris.",
        turns=[
            {
                "role": "user",
                "content": "I need a hotel in Paris near the Louvre."
            },
            {
                "role": "assistant",
                "content": "Hôtel du Louvre has rooms available. Which dates?"
            }
        ],
        context=["Hôtel du Louvre is a five-minute walk from the museum."],
        additional_metadata={"source": "faq"},
        comments="Reviewed by the support team.",
        source_file="capitals.csv",
        source_files=["capitals.csv"],
        finalized=True,
        custom_column_key_values={"difficulty": "easy"},
        images_mapping={
            "map": {"url": "https://example.com/paris.png", "local": False}
        },
        tags=["geography"]
    ),
)

For async mode, call a_update_golden and await it as shown below:

result = await client.datasets.a_update_golden(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
golden_idstrRequired. The unique id of the golden, returned when the dataset is pulled.
goldenGoldenRequestRequired. One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn. Pass a SingleTurnGoldenRequest or a MultiTurnGoldenRequest. See GoldenRequest.

Returns

This method returns an object of type GoldenRef.

Delete Golden

Permanently deletes a single golden. The rest of the dataset is unchanged, and this action cannot be undone.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.delete_golden(
    dataset_id="<DATASET-ID>",
    golden_id="<GOLDEN-ID>",
)

For async mode, call a_delete_golden and await it as shown below:

result = await client.datasets.a_delete_golden(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
golden_idstrRequired. The unique id of the golden, returned when the dataset is pulled.

Returns

This method returns an object of type GoldenRef.

Types

Golden

A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is.

Golden = Union[
    SingleTurnGolden,
    MultiTurnGolden,
]

A Golden is one of the shapes below. Send the fields of one of them, never a mix of both.

A single-turn golden as stored in the dataset.

class SingleTurnGolden:
    input: str
    actual_output: Optional[str] = Field(default=None, alias="actualOutput")
    expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
    context: Optional[List[str]] = None
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
    token_cost: Optional[float] = Field(default=None, alias="tokenCost")
    input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
    output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
    id: Optional[str] = None
    additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
    comments: Optional[str] = None
    source_file: Optional[str] = Field(default=None, alias="sourceFile")
    source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
    finalized: Optional[bool] = None
    custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
    tags: Optional[List[str]] = None

inputstrRequired

This is the input to your LLM application.

Example: "What is the capital of France?"

actual_outputOptional[str]

This is the actual output of your LLM application.

Example: "The capital of France is Paris."

expected_outputOptional[str]

This is the expected output of your LLM application, which is the ideal actual output.

Example: "Paris."

contextOptional[List[str]]

This is the ideal retrieval context of your LLM application.

Example: ["Paris is the capital of France."]

retrieval_contextOptional[List[str]]

This is the retrieval context of your LLM application.

Example: ["Paris is the capital and largest city of France."]

tools_calledOptional[List[ToolCall]]

This is the tools called by your LLM application.

See ToolCall.

expected_toolsOptional[List[ToolCall]]

This is the expected tools to be called by the LLM application.

See ToolCall.

token_costOptional[float]

This is the cost of the tokens used to produce the actual output.

Example: 0.002

input_token_countOptional[int]

This is the number of input tokens passed to the LLM model.

Example: 12

output_token_countOptional[int]

This is the number of output tokens generated by the LLM model.

Example: 3

idOptional[str]

The id of the golden assigned by Confident AI. Use it to get, update or delete this golden.

Example: "<GOLDEN-ID>"

additional_metadataOptional[Dict[str, Any]]

This is any additional metadata associated with the golden.

Example: {"source":"faq"}

commentsOptional[str]

This is any comments associated with the golden.

Example: "Reviewed by the support team."

source_fileOptional[str]

This is the source file from which the golden was retrieved.

Example: "capitals.csv"

source_filesOptional[List[str]]

These are the source files the golden was retrieved from.

Example: ["capitals.csv"]

finalizedOptional[bool]

This is true when the golden is finalized and ready to use in evaluations.

Example: true

custom_column_key_valuesOptional[Dict[str, str]]

Key-value pairs representing custom table column data for this golden. Keys correspond to the custom column keys defined in the dataset. Absent when the golden has no custom column values.

Example: {"difficulty":"easy"}

tagsOptional[List[str]]

These are the tags associated with the golden.

Example: ["geography"]

GoldenRef

class GoldenRef:
    id: str

idstrRequired

This is the unique id of the golden.

Example: "<GOLDEN-ID>"

GoldenRequest

One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn.

GoldenRequest = Union[
    SingleTurnGoldenRequest,
    MultiTurnGoldenRequest,
]

A GoldenRequest is one of the shapes below. Send the fields of one of them, never a mix of both.

A single-turn golden to write: one input to your LLM application and the outputs expected of it.

class SingleTurnGoldenRequest:
    input: str
    actual_output: Optional[str] = Field(default=None, alias="actualOutput")
    expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
    context: Optional[List[str]] = None
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
    token_cost: Optional[float] = Field(default=None, alias="tokenCost")
    input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
    output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
    additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
    comments: Optional[str] = None
    source_file: Optional[str] = Field(default=None, alias="sourceFile")
    source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
    finalized: Optional[bool] = None
    custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
    images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
    tags: Optional[List[str]] = None

inputstrRequired

This is the input to your LLM application.

Example: "What is the capital of France?"

actual_outputOptional[str]

This is the actual output of your LLM application.

Example: "The capital of France is Paris."

expected_outputOptional[str]

This is the expected output of your LLM application, which is the ideal actual output.

Example: "Paris."

contextOptional[List[str]]

This is the ideal retrieval context of your LLM application.

Example: ["Paris is the capital of France."]

retrieval_contextOptional[List[str]]

This is the retrieval context of your LLM application.

Example: ["Paris is the capital and largest city of France."]

tools_calledOptional[List[ToolCall]]

This is the tools called by your LLM application.

See ToolCall.

expected_toolsOptional[List[ToolCall]]

This is the expected tools to be called by the LLM application.

See ToolCall.

token_costOptional[float]

This is the cost of the tokens used to produce the actual output.

Example: 0.002

input_token_countOptional[int]

This is the number of input tokens passed to the LLM model.

Example: 12

output_token_countOptional[int]

This is the number of output tokens generated by the LLM model.

Example: 3

additional_metadataOptional[Dict[str, Any]]

Additional metadata to associate with the golden.

Example: {"source":"faq"}

commentsOptional[str]

Comments to associate with the golden.

Example: "Reviewed by the support team."

source_fileOptional[str]

The source file the golden was retrieved from. Like tags and customColumnKeyValues, this is left unchanged when the request omits it; send null to clear it.

Example: "capitals.csv"

source_filesOptional[List[str]]

The source files the golden was retrieved from. Like tags and customColumnKeyValues, these are left unchanged when the request omits them; send an empty array to clear them.

Example: ["capitals.csv"]

finalizedOptional[bool]

Whether the golden is ready to use in evaluations. When pushing or queueing a list of goldens the request decides this for every golden and this field is ignored.

Example: true

custom_column_key_valuesOptional[Dict[str, str]]

Custom dataset column values keyed by column name. A column that does not exist in the dataset yet is created.

Example: {"difficulty":"easy"}

images_mappingOptional[Dict[str, MLLMImage]]

The media this golden refers to, keyed by the id inside each placeholder. Put [DEEPEVAL:IMAGE:<id>] or [DEEPEVAL:PDF:<id>] in a text field where the media belongs, and the platform substitutes the entry with a matching key.

See MLLMImage.

Example: {"map":{"url":"https://example.com/paris.png","local":false}}

tagsOptional[List[str]]

Tags to associate with the golden, which is useful for grouping and filtering goldens. A tag that does not exist in the dataset yet is created.

Example: ["geography"]

MLLMImage

An image referenced from a text field by a [DEEPEVAL:IMAGE:<key>] marker. Send either a public url or the bytes in base64.

class MLLMImage:
    url: str
    local: bool
    base64: Optional[str] = None
    filename: Optional[str] = None
    mime_type: Optional[str] = Field(default=None, alias="mimeType")
    data_base64: Optional[str] = Field(default=None, alias="dataBase64")

urlstrRequired

This is the URL of the image.

Example: "https://example.com/everest.png"

localboolRequired

This is true when the image is your local file.

Example: false

base64Optional[str]

The base64 data of the image.

Example: "iVBORw0KGgo="

filenameOptional[str]

The original file name.

Example: "everest.png"

mime_typeOptional[str]

The image's MIME type.

Example: "image/png"

data_base64Optional[str]

The image encoded as a base64 data URL.

Example: "data:image/png;base64,iVBORw0KGgo="

ToolCall

A tool your LLM application invoked, with what it passed in and what came back.

class ToolCall:
    name: str
    type: Optional[ToolCallType] = None
    description: Optional[str] = None
    input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
    output: Optional[Any] = None
    reasoning: Optional[str] = None

namestrRequired

This is the name of the tool.

Example: "get_landmark_info"

typeOptional[ToolCallType]

descriptionOptional[str]

This is the description of the tool.

Example: "This tool gives information about a mountain."

input_parametersOptional[Dict[str, Any]]

This is the input parameters that are passed to the tool.

Example: {"mountain":"Everest"}

outputOptional[Any]

This is the output of the tool.

Example: "8,848 metres"

reasoningOptional[str]

This is the reasoning your LLM provided for the tool call.

Example: "The user asked for the height of a mountain."

ToolCallType

The type of the tool call, either a function or an MCP tool.

class ToolCallType(Enum):
    FUNCTION = "FUNCTION"
    MCP = "MCP"

FUNCTION · MCP

Turn

One message in a conversation, from either the user or the assistant, with the context and tools behind an assistant reply.

class Turn:
    id: Optional[str] = None
    role: TurnRole
    content: str
    user_id: Optional[str] = Field(default=None, alias="userId")
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")

idOptional[str]

The id of a turn assigned by Confident AI.

Example: "<TURN-ID>"

roleTurnRoleRequired

contentstrRequired

The message content of the turn.

Example: "How tall is Mount Everest?"

user_idOptional[str]

The user ID associated with the turn.

Example: "end-user-42"

retrieval_contextOptional[List[str]]

The contexts retrieved to generate the LLM response for this turn.

Example: ["Everest is 8,848 metres tall."]

tools_calledOptional[List[ToolCall]]

The tools called to generate the LLM response for this turn.

See ToolCall.

TurnRole

The role of the turn, either user or assistant.

class TurnRole(Enum):
    USER = "user"
    ASSISTANT = "assistant"

USER · ASSISTANT

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI