Goldens
Overview
The Confident AI SDK exposes every Golden method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
Single Turn Golden
Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.
A single-turn golden to write: one input to your LLM application and the outputs expected of it.
from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
client = ConfidentAI()
dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.create_golden(
golden=SingleTurnGoldenRequest(
input="What is the capital of France?",
actual_output="The capital of France is Paris.",
expected_output="Paris.",
context=["Paris is the capital of France."],
retrieval_context=[
"Paris is the capital and largest city of France."
],
tools_called=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
expected_tools=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
token_cost=0.002,
input_token_count=12,
output_token_count=3,
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
images_mapping={
"map": {"url": "https://example.com/paris.png", "local": False}
},
tags=["geography"]
),
version="00.00.01",
)For async mode, call a_create_golden and await it as shown below:
result = await dataset.a_create_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
golden | GoldenRequest | Required. See GoldenRequest. |
version | Optional[str] | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
import { ConfidentAI } from "confident-ai";
import { ToolCallType } from "confident-ai/common";
const client = new ConfidentAI();
const dataset = client.dataset("<DATASET-ID>");
const result = await dataset.createGolden(
{
input: "What is the capital of France?",
actualOutput: "The capital of France is Paris.",
expectedOutput: "Paris.",
context: ["Paris is the capital of France."],
retrievalContext: ["Paris is the capital and largest city of France."],
toolsCalled: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
expectedTools: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
tokenCost: 0.002,
inputTokenCount: 12,
outputTokenCount: 3,
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
imagesMapping: {
map: { url: "https://example.com/paris.png", local: false }
},
tags: ["geography"]
},
{ version: "00.00.01" },
);Parameters
| Parameter | Type | Description |
|---|---|---|
golden | GoldenRequest | Required. See GoldenRequest. |
version | string | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
Returns
This method returns an object of type GoldenRef.
Multi Turn Golden
Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.
A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.
from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest
client = ConfidentAI()
dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.create_golden(
golden=MultiTurnGoldenRequest(
scenario="A traveller wants to book a hotel in Paris.",
expected_outcome="The assistant confirms a reservation near the Louvre.",
user_description="A traveller planning a weekend in Paris.",
turns=[
{
"role": "user",
"content": "I need a hotel in Paris near the Louvre."
},
{
"role": "assistant",
"content": "Hôtel du Louvre has rooms available. Which dates?"
}
],
context=["Hôtel du Louvre is a five-minute walk from the museum."],
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
images_mapping={
"map": {"url": "https://example.com/paris.png", "local": False}
},
tags=["geography"]
),
version="00.00.01",
)For async mode, call a_create_golden and await it as shown below:
result = await dataset.a_create_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
golden | GoldenRequest | Required. See GoldenRequest. |
version | Optional[str] | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const dataset = client.dataset("<DATASET-ID>");
const result = await dataset.createGolden(
{
scenario: "A traveller wants to book a hotel in Paris.",
expectedOutcome: "The assistant confirms a reservation near the Louvre.",
userDescription: "A traveller planning a weekend in Paris.",
turns: [
{
role: "user",
content: "I need a hotel in Paris near the Louvre."
},
{
role: "assistant",
content: "Hôtel du Louvre has rooms available. Which dates?"
}
],
context: ["Hôtel du Louvre is a five-minute walk from the museum."],
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
imagesMapping: {
map: { url: "https://example.com/paris.png", local: false }
},
tags: ["geography"]
},
{ version: "00.00.01" },
);Parameters
| Parameter | Type | Description |
|---|---|---|
golden | GoldenRequest | Required. See GoldenRequest. |
version | string | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
Returns
This method returns an object of type GoldenRef.
Get Golden
Retrieves a single golden by id. It is single-turn or multi-turn according to the dataset's multiTurn.
from confident_ai import ConfidentAI
client = ConfidentAI()
dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.get_golden(golden_id="<GOLDEN-ID>")For async mode, call a_get_golden and await it as shown below:
result = await dataset.a_get_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
golden_id | str | Required. The unique id of the golden, returned when the dataset is pulled. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const dataset = client.dataset("<DATASET-ID>");
const result = await dataset.getGolden("<GOLDEN-ID>");Parameters
| Parameter | Type | Description |
|---|---|---|
goldenId | string | Required. The unique id of the golden, returned when the dataset is pulled. |
Returns
This method returns an object of type Golden.
Delete Golden
Permanently deletes a single golden. The rest of the dataset is unchanged, and this action cannot be undone.
from confident_ai import ConfidentAI
client = ConfidentAI()
dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.delete_golden(golden_id="<GOLDEN-ID>")For async mode, call a_delete_golden and await it as shown below:
result = await dataset.a_delete_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
golden_id | str | Required. The unique id of the golden, returned when the dataset is pulled. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const dataset = client.dataset("<DATASET-ID>");
const result = await dataset.deleteGolden("<GOLDEN-ID>");Parameters
| Parameter | Type | Description |
|---|---|---|
goldenId | string | Required. The unique id of the golden, returned when the dataset is pulled. |
Returns
This method returns an object of type GoldenRef.
Single Turn Golden
Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.
A single-turn golden as stored in the dataset.
from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGolden
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
client = ConfidentAI()
dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.update_golden(
golden=SingleTurnGolden(
input="What is the capital of France?",
actual_output="The capital of France is Paris.",
expected_output="Paris.",
context=["Paris is the capital of France."],
retrieval_context=[
"Paris is the capital and largest city of France."
],
tools_called=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
expected_tools=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
token_cost=0.002,
input_token_count=12,
output_token_count=3,
id="<GOLDEN-ID>",
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
tags=["geography"]
),
)For async mode, call a_update_golden and await it as shown below:
result = await dataset.a_update_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
golden | Golden | Required. A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is. See GoldenRequest. |
import { ConfidentAI } from "confident-ai";
import { ToolCallType } from "confident-ai/common";
const client = new ConfidentAI();
const dataset = client.dataset("<DATASET-ID>");
const result = await dataset.updateGolden(
{
input: "What is the capital of France?",
actualOutput: "The capital of France is Paris.",
expectedOutput: "Paris.",
context: ["Paris is the capital of France."],
retrievalContext: ["Paris is the capital and largest city of France."],
toolsCalled: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
expectedTools: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
tokenCost: 0.002,
inputTokenCount: 12,
outputTokenCount: 3,
id: "<GOLDEN-ID>",
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
tags: ["geography"]
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
golden | Golden | Required. A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is. See GoldenRequest. |
Returns
This method returns an object of type GoldenRef.
Multi Turn Golden
Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.
A multi-turn golden as stored in the dataset.
from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGolden
client = ConfidentAI()
dataset = client.dataset(dataset_id="<DATASET-ID>")
result = dataset.update_golden(
golden=MultiTurnGolden(
scenario="A traveller wants to book a hotel in Paris.",
expected_outcome="The assistant confirms a reservation near the Louvre.",
user_description="A traveller planning a weekend in Paris.",
turns=[
{
"role": "user",
"content": "I need a hotel in Paris near the Louvre."
},
{
"role": "assistant",
"content": "Hôtel du Louvre has rooms available. Which dates?"
}
],
context=["Hôtel du Louvre is a five-minute walk from the museum."],
id="<GOLDEN-ID>",
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
tags=["geography"]
),
)For async mode, call a_update_golden and await it as shown below:
result = await dataset.a_update_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
golden | Golden | Required. A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is. See GoldenRequest. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const dataset = client.dataset("<DATASET-ID>");
const result = await dataset.updateGolden(
{
scenario: "A traveller wants to book a hotel in Paris.",
expectedOutcome: "The assistant confirms a reservation near the Louvre.",
userDescription: "A traveller planning a weekend in Paris.",
turns: [
{
role: "user",
content: "I need a hotel in Paris near the Louvre."
},
{
role: "assistant",
content: "Hôtel du Louvre has rooms available. Which dates?"
}
],
context: ["Hôtel du Louvre is a five-minute walk from the museum."],
id: "<GOLDEN-ID>",
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
tags: ["geography"]
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
golden | Golden | Required. A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is. See GoldenRequest. |
Returns
This method returns an object of type GoldenRef.
Methods (Stateless)
These methods take every argument themselves, so a caller reaches them through client.datasets without opening a Dataset first.
Single Turn Golden
Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.
A single-turn golden to write: one input to your LLM application and the outputs expected of it.
from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
client = ConfidentAI()
result = client.datasets.create_golden(
dataset_id="<DATASET-ID>",
golden=SingleTurnGoldenRequest(
input="What is the capital of France?",
actual_output="The capital of France is Paris.",
expected_output="Paris.",
context=["Paris is the capital of France."],
retrieval_context=[
"Paris is the capital and largest city of France."
],
tools_called=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
expected_tools=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
token_cost=0.002,
input_token_count=12,
output_token_count=3,
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
images_mapping={
"map": {"url": "https://example.com/paris.png", "local": False}
},
tags=["geography"]
),
version="00.00.01",
)For async mode, call a_create_golden and await it as shown below:
result = await client.datasets.a_create_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
dataset_id | str | Required. The unique id of the dataset. |
golden | GoldenRequest | Required. See GoldenRequest. |
version | Optional[str] | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
import { ConfidentAI } from "confident-ai";
import { ToolCallType } from "confident-ai/common";
const client = new ConfidentAI();
const result = await client.datasets.createGolden(
"<DATASET-ID>",
{
input: "What is the capital of France?",
actualOutput: "The capital of France is Paris.",
expectedOutput: "Paris.",
context: ["Paris is the capital of France."],
retrievalContext: ["Paris is the capital and largest city of France."],
toolsCalled: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
expectedTools: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
tokenCost: 0.002,
inputTokenCount: 12,
outputTokenCount: 3,
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
imagesMapping: {
map: { url: "https://example.com/paris.png", local: false }
},
tags: ["geography"]
},
{ version: "00.00.01" },
);Parameters
| Parameter | Type | Description |
|---|---|---|
datasetId | string | Required. The unique id of the dataset. |
golden | GoldenRequest | Required. See GoldenRequest. |
version | string | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
Returns
This method returns an object of type GoldenRef.
Multi Turn Golden
Adds a single golden to the dataset and returns its id. Pass version to add it to a specific dataset version; omitting it targets the latest.
A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.
from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest
client = ConfidentAI()
result = client.datasets.create_golden(
dataset_id="<DATASET-ID>",
golden=MultiTurnGoldenRequest(
scenario="A traveller wants to book a hotel in Paris.",
expected_outcome="The assistant confirms a reservation near the Louvre.",
user_description="A traveller planning a weekend in Paris.",
turns=[
{
"role": "user",
"content": "I need a hotel in Paris near the Louvre."
},
{
"role": "assistant",
"content": "Hôtel du Louvre has rooms available. Which dates?"
}
],
context=["Hôtel du Louvre is a five-minute walk from the museum."],
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
images_mapping={
"map": {"url": "https://example.com/paris.png", "local": False}
},
tags=["geography"]
),
version="00.00.01",
)For async mode, call a_create_golden and await it as shown below:
result = await client.datasets.a_create_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
dataset_id | str | Required. The unique id of the dataset. |
golden | GoldenRequest | Required. See GoldenRequest. |
version | Optional[str] | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.datasets.createGolden(
"<DATASET-ID>",
{
scenario: "A traveller wants to book a hotel in Paris.",
expectedOutcome: "The assistant confirms a reservation near the Louvre.",
userDescription: "A traveller planning a weekend in Paris.",
turns: [
{
role: "user",
content: "I need a hotel in Paris near the Louvre."
},
{
role: "assistant",
content: "Hôtel du Louvre has rooms available. Which dates?"
}
],
context: ["Hôtel du Louvre is a five-minute walk from the museum."],
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
imagesMapping: {
map: { url: "https://example.com/paris.png", local: false }
},
tags: ["geography"]
},
{ version: "00.00.01" },
);Parameters
| Parameter | Type | Description |
|---|---|---|
datasetId | string | Required. The unique id of the dataset. |
golden | GoldenRequest | Required. See GoldenRequest. |
version | string | The dataset version to add the golden to. Omitting it targets the latest version, or the unversioned goldens when the dataset has no versions. |
Returns
This method returns an object of type GoldenRef.
Get Golden
Retrieves a single golden by id. It is single-turn or multi-turn according to the dataset's multiTurn.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.datasets.get_golden(
dataset_id="<DATASET-ID>",
golden_id="<GOLDEN-ID>",
)For async mode, call a_get_golden and await it as shown below:
result = await client.datasets.a_get_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
dataset_id | str | Required. The unique id of the dataset. |
golden_id | str | Required. The unique id of the golden, returned when the dataset is pulled. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.datasets.getGolden(
"<DATASET-ID>",
"<GOLDEN-ID>",
);Parameters
| Parameter | Type | Description |
|---|---|---|
datasetId | string | Required. The unique id of the dataset. |
goldenId | string | Required. The unique id of the golden, returned when the dataset is pulled. |
Returns
This method returns an object of type Golden.
Single Turn Golden
Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.
A single-turn golden to write: one input to your LLM application and the outputs expected of it.
from confident_ai import ConfidentAI
from confident_ai.datasets import SingleTurnGoldenRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
client = ConfidentAI()
result = client.datasets.update_golden(
dataset_id="<DATASET-ID>",
golden_id="<GOLDEN-ID>",
golden=SingleTurnGoldenRequest(
input="What is the capital of France?",
actual_output="The capital of France is Paris.",
expected_output="Paris.",
context=["Paris is the capital of France."],
retrieval_context=[
"Paris is the capital and largest city of France."
],
tools_called=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
expected_tools=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
token_cost=0.002,
input_token_count=12,
output_token_count=3,
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
images_mapping={
"map": {"url": "https://example.com/paris.png", "local": False}
},
tags=["geography"]
),
)For async mode, call a_update_golden and await it as shown below:
result = await client.datasets.a_update_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
dataset_id | str | Required. The unique id of the dataset. |
golden_id | str | Required. The unique id of the golden, returned when the dataset is pulled. |
golden | GoldenRequest | Required. One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn. Pass a SingleTurnGoldenRequest or a MultiTurnGoldenRequest. See GoldenRequest. |
import { ConfidentAI } from "confident-ai";
import { ToolCallType } from "confident-ai/common";
const client = new ConfidentAI();
const result = await client.datasets.updateGolden(
"<DATASET-ID>",
"<GOLDEN-ID>",
{
input: "What is the capital of France?",
actualOutput: "The capital of France is Paris.",
expectedOutput: "Paris.",
context: ["Paris is the capital of France."],
retrievalContext: ["Paris is the capital and largest city of France."],
toolsCalled: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
expectedTools: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
tokenCost: 0.002,
inputTokenCount: 12,
outputTokenCount: 3,
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
imagesMapping: {
map: { url: "https://example.com/paris.png", local: false }
},
tags: ["geography"]
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
datasetId | string | Required. The unique id of the dataset. |
goldenId | string | Required. The unique id of the golden, returned when the dataset is pulled. |
golden | GoldenRequest | Required. One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn. Pass a SingleTurnGoldenRequest or a MultiTurnGoldenRequest. See GoldenRequest. |
Returns
This method returns an object of type GoldenRef.
Multi Turn Golden
Replaces the fields of a single golden with the values you send and returns its id. sourceFile, sourceFiles, tags and customColumnKeyValues are left unchanged unless you include them. The golden's kind must match the dataset's multiTurn.
A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.
from confident_ai import ConfidentAI
from confident_ai.datasets import MultiTurnGoldenRequest
client = ConfidentAI()
result = client.datasets.update_golden(
dataset_id="<DATASET-ID>",
golden_id="<GOLDEN-ID>",
golden=MultiTurnGoldenRequest(
scenario="A traveller wants to book a hotel in Paris.",
expected_outcome="The assistant confirms a reservation near the Louvre.",
user_description="A traveller planning a weekend in Paris.",
turns=[
{
"role": "user",
"content": "I need a hotel in Paris near the Louvre."
},
{
"role": "assistant",
"content": "Hôtel du Louvre has rooms available. Which dates?"
}
],
context=["Hôtel du Louvre is a five-minute walk from the museum."],
additional_metadata={"source": "faq"},
comments="Reviewed by the support team.",
source_file="capitals.csv",
source_files=["capitals.csv"],
finalized=True,
custom_column_key_values={"difficulty": "easy"},
images_mapping={
"map": {"url": "https://example.com/paris.png", "local": False}
},
tags=["geography"]
),
)For async mode, call a_update_golden and await it as shown below:
result = await client.datasets.a_update_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
dataset_id | str | Required. The unique id of the dataset. |
golden_id | str | Required. The unique id of the golden, returned when the dataset is pulled. |
golden | GoldenRequest | Required. One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn. Pass a SingleTurnGoldenRequest or a MultiTurnGoldenRequest. See GoldenRequest. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.datasets.updateGolden(
"<DATASET-ID>",
"<GOLDEN-ID>",
{
scenario: "A traveller wants to book a hotel in Paris.",
expectedOutcome: "The assistant confirms a reservation near the Louvre.",
userDescription: "A traveller planning a weekend in Paris.",
turns: [
{
role: "user",
content: "I need a hotel in Paris near the Louvre."
},
{
role: "assistant",
content: "Hôtel du Louvre has rooms available. Which dates?"
}
],
context: ["Hôtel du Louvre is a five-minute walk from the museum."],
additionalMetadata: { source: "faq" },
comments: "Reviewed by the support team.",
sourceFile: "capitals.csv",
sourceFiles: ["capitals.csv"],
finalized: true,
customColumnKeyValues: { difficulty: "easy" },
imagesMapping: {
map: { url: "https://example.com/paris.png", local: false }
},
tags: ["geography"]
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
datasetId | string | Required. The unique id of the dataset. |
goldenId | string | Required. The unique id of the golden, returned when the dataset is pulled. |
golden | GoldenRequest | Required. One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn. Pass a SingleTurnGoldenRequest or a MultiTurnGoldenRequest. See GoldenRequest. |
Returns
This method returns an object of type GoldenRef.
Delete Golden
Permanently deletes a single golden. The rest of the dataset is unchanged, and this action cannot be undone.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.datasets.delete_golden(
dataset_id="<DATASET-ID>",
golden_id="<GOLDEN-ID>",
)For async mode, call a_delete_golden and await it as shown below:
result = await client.datasets.a_delete_golden(...)Parameters
| Parameter | Type | Description |
|---|---|---|
dataset_id | str | Required. The unique id of the dataset. |
golden_id | str | Required. The unique id of the golden, returned when the dataset is pulled. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.datasets.deleteGolden(
"<DATASET-ID>",
"<GOLDEN-ID>",
);Parameters
| Parameter | Type | Description |
|---|---|---|
datasetId | string | Required. The unique id of the dataset. |
goldenId | string | Required. The unique id of the golden, returned when the dataset is pulled. |
Returns
This method returns an object of type GoldenRef.
Types
Golden
A golden in the dataset: single-turn when it carries input, multi-turn when it carries scenario. The dataset's multiTurn decides which kind every golden in it is.
Golden = Union[
SingleTurnGolden,
MultiTurnGolden,
]type Golden =
| SingleTurnGolden
| MultiTurnGolden;A Golden is one of the shapes below. Send the fields of one of them, never a mix of both.
A single-turn golden as stored in the dataset.
class SingleTurnGolden:
input: str
actual_output: Optional[str] = Field(default=None, alias="actualOutput")
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
context: Optional[List[str]] = None
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
token_cost: Optional[float] = Field(default=None, alias="tokenCost")
input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
id: Optional[str] = None
additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
comments: Optional[str] = None
source_file: Optional[str] = Field(default=None, alias="sourceFile")
source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
finalized: Optional[bool] = None
custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
tags: Optional[List[str]] = NoneinputstrRequired
This is the input to your LLM application.
Example: "What is the capital of France?"
actual_outputOptional[str]
This is the actual output of your LLM application.
Example: "The capital of France is Paris."
expected_outputOptional[str]
This is the expected output of your LLM application, which is the ideal actual output.
Example: "Paris."
contextOptional[List[str]]
This is the ideal retrieval context of your LLM application.
Example: ["Paris is the capital of France."]
retrieval_contextOptional[List[str]]
This is the retrieval context of your LLM application.
Example: ["Paris is the capital and largest city of France."]
tools_calledOptional[List[ToolCall]]
This is the tools called by your LLM application.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the LLM application.
See ToolCall.
token_costOptional[float]
This is the cost of the tokens used to produce the actual output.
Example: 0.002
input_token_countOptional[int]
This is the number of input tokens passed to the LLM model.
Example: 12
output_token_countOptional[int]
This is the number of output tokens generated by the LLM model.
Example: 3
idOptional[str]
The id of the golden assigned by Confident AI. Use it to get, update or delete this golden.
Example: "<GOLDEN-ID>"
additional_metadataOptional[Dict[str, Any]]
This is any additional metadata associated with the golden.
Example: {"source":"faq"}
commentsOptional[str]
This is any comments associated with the golden.
Example: "Reviewed by the support team."
source_fileOptional[str]
This is the source file from which the golden was retrieved.
Example: "capitals.csv"
source_filesOptional[List[str]]
These are the source files the golden was retrieved from.
Example: ["capitals.csv"]
finalizedOptional[bool]
This is true when the golden is finalized and ready to use in evaluations.
Example: true
custom_column_key_valuesOptional[Dict[str, str]]
Key-value pairs representing custom table column data for this golden. Keys correspond to the custom column keys defined in the dataset. Absent when the golden has no custom column values.
Example: {"difficulty":"easy"}
tagsOptional[List[str]]
These are the tags associated with the golden.
Example: ["geography"]
interface SingleTurnGolden {
input: string;
actualOutput?: string | null;
expectedOutput?: string | null;
context?: string[] | null;
retrievalContext?: string[] | null;
toolsCalled?: ToolCall[] | null;
expectedTools?: ToolCall[] | null;
tokenCost?: number | null;
inputTokenCount?: number | null;
outputTokenCount?: number | null;
id?: string;
additionalMetadata?: Record<string, unknown> | null;
comments?: string | null;
sourceFile?: string | null;
sourceFiles?: string[];
finalized?: boolean;
customColumnKeyValues?: Record<string, string>;
tags?: string[];
}inputstringRequired
This is the input to your LLM application.
Example: "What is the capital of France?"
actualOutputstring | null
This is the actual output of your LLM application.
Example: "The capital of France is Paris."
expectedOutputstring | null
This is the expected output of your LLM application, which is the ideal actual output.
Example: "Paris."
contextstring[] | null
This is the ideal retrieval context of your LLM application.
Example: ["Paris is the capital of France."]
retrievalContextstring[] | null
This is the retrieval context of your LLM application.
Example: ["Paris is the capital and largest city of France."]
toolsCalledToolCall[] | null
This is the tools called by your LLM application.
See ToolCall.
expectedToolsToolCall[] | null
This is the expected tools to be called by the LLM application.
See ToolCall.
tokenCostnumber | null
This is the cost of the tokens used to produce the actual output.
Example: 0.002
inputTokenCountnumber | null
This is the number of input tokens passed to the LLM model.
Example: 12
outputTokenCountnumber | null
This is the number of output tokens generated by the LLM model.
Example: 3
idstring
The id of the golden assigned by Confident AI. Use it to get, update or delete this golden.
Example: "<GOLDEN-ID>"
additionalMetadataRecord<string, unknown> | null
This is any additional metadata associated with the golden.
Example: {"source":"faq"}
commentsstring | null
This is any comments associated with the golden.
Example: "Reviewed by the support team."
sourceFilestring | null
This is the source file from which the golden was retrieved.
Example: "capitals.csv"
sourceFilesstring[]
These are the source files the golden was retrieved from.
Example: ["capitals.csv"]
finalizedboolean
This is true when the golden is finalized and ready to use in evaluations.
Example: true
customColumnKeyValuesRecord<string, string>
Key-value pairs representing custom table column data for this golden. Keys correspond to the custom column keys defined in the dataset. Absent when the golden has no custom column values.
Example: {"difficulty":"easy"}
tagsstring[]
These are the tags associated with the golden.
Example: ["geography"]
A multi-turn golden as stored in the dataset.
class MultiTurnGolden:
scenario: str
expected_outcome: Optional[str] = Field(default=None, alias="expectedOutcome")
user_description: Optional[str] = Field(default=None, alias="userDescription")
turns: Optional[List[Turn]] = None
context: Optional[List[str]] = None
id: Optional[str] = None
additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
comments: Optional[str] = None
source_file: Optional[str] = Field(default=None, alias="sourceFile")
source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
finalized: Optional[bool] = None
custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
tags: Optional[List[str]] = NonescenariostrRequired
This is a description of the conversation context.
Example: "A traveller wants to book a hotel in Paris."
expected_outcomeOptional[str]
This describes the expected outcome, or ideal conversation flow, of the conversation.
Example: "The assistant confirms a reservation near the Louvre."
user_descriptionOptional[str]
This is the description of the user in the conversation.
Example: "A traveller planning a weekend in Paris."
turnsOptional[List[Turn]]
This is the list of turns in the conversation.
See Turn.
Example: [{"role":"user","content":"I need a hotel in Paris near the Louvre."},{"role":"assistant","content":"Hôtel du Louvre has rooms available. Which dates?"}]
contextOptional[List[str]]
This is the context of the conversation.
Example: ["Hôtel du Louvre is a five-minute walk from the museum."]
idOptional[str]
The id of the golden assigned by Confident AI. Use it to get, update or delete this golden.
Example: "<GOLDEN-ID>"
additional_metadataOptional[Dict[str, Any]]
This is any additional metadata associated with the golden.
Example: {"source":"faq"}
commentsOptional[str]
This is any comments associated with the golden.
Example: "Reviewed by the support team."
source_fileOptional[str]
This is the source file from which the golden was retrieved.
Example: "capitals.csv"
source_filesOptional[List[str]]
These are the source files the golden was retrieved from.
Example: ["capitals.csv"]
finalizedOptional[bool]
This is true when the golden is finalized and ready to use in evaluations.
Example: true
custom_column_key_valuesOptional[Dict[str, str]]
Key-value pairs representing custom table column data for this golden. Keys correspond to the custom column keys defined in the dataset. Absent when the golden has no custom column values.
Example: {"difficulty":"easy"}
tagsOptional[List[str]]
These are the tags associated with the golden.
Example: ["geography"]
interface MultiTurnGolden {
scenario: string;
expectedOutcome?: string | null;
userDescription?: string | null;
turns?: Turn[] | null;
context?: string[] | null;
id?: string;
additionalMetadata?: Record<string, unknown> | null;
comments?: string | null;
sourceFile?: string | null;
sourceFiles?: string[];
finalized?: boolean;
customColumnKeyValues?: Record<string, string>;
tags?: string[];
}scenariostringRequired
This is a description of the conversation context.
Example: "A traveller wants to book a hotel in Paris."
expectedOutcomestring | null
This describes the expected outcome, or ideal conversation flow, of the conversation.
Example: "The assistant confirms a reservation near the Louvre."
userDescriptionstring | null
This is the description of the user in the conversation.
Example: "A traveller planning a weekend in Paris."
turnsTurn[] | null
This is the list of turns in the conversation.
See Turn.
Example: [{"role":"user","content":"I need a hotel in Paris near the Louvre."},{"role":"assistant","content":"Hôtel du Louvre has rooms available. Which dates?"}]
contextstring[] | null
This is the context of the conversation.
Example: ["Hôtel du Louvre is a five-minute walk from the museum."]
idstring
The id of the golden assigned by Confident AI. Use it to get, update or delete this golden.
Example: "<GOLDEN-ID>"
additionalMetadataRecord<string, unknown> | null
This is any additional metadata associated with the golden.
Example: {"source":"faq"}
commentsstring | null
This is any comments associated with the golden.
Example: "Reviewed by the support team."
sourceFilestring | null
This is the source file from which the golden was retrieved.
Example: "capitals.csv"
sourceFilesstring[]
These are the source files the golden was retrieved from.
Example: ["capitals.csv"]
finalizedboolean
This is true when the golden is finalized and ready to use in evaluations.
Example: true
customColumnKeyValuesRecord<string, string>
Key-value pairs representing custom table column data for this golden. Keys correspond to the custom column keys defined in the dataset. Absent when the golden has no custom column values.
Example: {"difficulty":"easy"}
tagsstring[]
These are the tags associated with the golden.
Example: ["geography"]
GoldenRef
class GoldenRef:
id: stridstrRequired
This is the unique id of the golden.
Example: "<GOLDEN-ID>"
interface GoldenRef {
id: string;
}idstringRequired
This is the unique id of the golden.
Example: "<GOLDEN-ID>"
GoldenRequest
One golden to write: single-turn when it carries input, multi-turn when it carries scenario. A golden cannot be both, and its kind must match the dataset's multiTurn.
GoldenRequest = Union[
SingleTurnGoldenRequest,
MultiTurnGoldenRequest,
]type GoldenRequest =
| SingleTurnGoldenRequest
| MultiTurnGoldenRequest;A GoldenRequest is one of the shapes below. Send the fields of one of them, never a mix of both.
A single-turn golden to write: one input to your LLM application and the outputs expected of it.
class SingleTurnGoldenRequest:
input: str
actual_output: Optional[str] = Field(default=None, alias="actualOutput")
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
context: Optional[List[str]] = None
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
token_cost: Optional[float] = Field(default=None, alias="tokenCost")
input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
comments: Optional[str] = None
source_file: Optional[str] = Field(default=None, alias="sourceFile")
source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
finalized: Optional[bool] = None
custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
tags: Optional[List[str]] = NoneinputstrRequired
This is the input to your LLM application.
Example: "What is the capital of France?"
actual_outputOptional[str]
This is the actual output of your LLM application.
Example: "The capital of France is Paris."
expected_outputOptional[str]
This is the expected output of your LLM application, which is the ideal actual output.
Example: "Paris."
contextOptional[List[str]]
This is the ideal retrieval context of your LLM application.
Example: ["Paris is the capital of France."]
retrieval_contextOptional[List[str]]
This is the retrieval context of your LLM application.
Example: ["Paris is the capital and largest city of France."]
tools_calledOptional[List[ToolCall]]
This is the tools called by your LLM application.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the LLM application.
See ToolCall.
token_costOptional[float]
This is the cost of the tokens used to produce the actual output.
Example: 0.002
input_token_countOptional[int]
This is the number of input tokens passed to the LLM model.
Example: 12
output_token_countOptional[int]
This is the number of output tokens generated by the LLM model.
Example: 3
additional_metadataOptional[Dict[str, Any]]
Additional metadata to associate with the golden.
Example: {"source":"faq"}
commentsOptional[str]
Comments to associate with the golden.
Example: "Reviewed by the support team."
source_fileOptional[str]
The source file the golden was retrieved from. Like tags and customColumnKeyValues, this is left unchanged when the request omits it; send null to clear it.
Example: "capitals.csv"
source_filesOptional[List[str]]
The source files the golden was retrieved from. Like tags and customColumnKeyValues, these are left unchanged when the request omits them; send an empty array to clear them.
Example: ["capitals.csv"]
finalizedOptional[bool]
Whether the golden is ready to use in evaluations. When pushing or queueing a list of goldens the request decides this for every golden and this field is ignored.
Example: true
custom_column_key_valuesOptional[Dict[str, str]]
Custom dataset column values keyed by column name. A column that does not exist in the dataset yet is created.
Example: {"difficulty":"easy"}
images_mappingOptional[Dict[str, MLLMImage]]
The media this golden refers to, keyed by the id inside each placeholder. Put [DEEPEVAL:IMAGE:<id>] or [DEEPEVAL:PDF:<id>] in a text field where the media belongs, and the platform substitutes the entry with a matching key.
See MLLMImage.
Example: {"map":{"url":"https://example.com/paris.png","local":false}}
tagsOptional[List[str]]
Tags to associate with the golden, which is useful for grouping and filtering goldens. A tag that does not exist in the dataset yet is created.
Example: ["geography"]
interface SingleTurnGoldenRequest {
input: string;
actualOutput?: string | null;
expectedOutput?: string | null;
context?: string[] | null;
retrievalContext?: string[] | null;
toolsCalled?: ToolCall[] | null;
expectedTools?: ToolCall[] | null;
tokenCost?: number | null;
inputTokenCount?: number | null;
outputTokenCount?: number | null;
additionalMetadata?: Record<string, unknown> | null;
comments?: string | null;
sourceFile?: string | null;
sourceFiles?: string[];
finalized?: boolean;
customColumnKeyValues?: Record<string, string>;
imagesMapping?: Record<string, MLLMImage>;
tags?: string[];
}inputstringRequired
This is the input to your LLM application.
Example: "What is the capital of France?"
actualOutputstring | null
This is the actual output of your LLM application.
Example: "The capital of France is Paris."
expectedOutputstring | null
This is the expected output of your LLM application, which is the ideal actual output.
Example: "Paris."
contextstring[] | null
This is the ideal retrieval context of your LLM application.
Example: ["Paris is the capital of France."]
retrievalContextstring[] | null
This is the retrieval context of your LLM application.
Example: ["Paris is the capital and largest city of France."]
toolsCalledToolCall[] | null
This is the tools called by your LLM application.
See ToolCall.
expectedToolsToolCall[] | null
This is the expected tools to be called by the LLM application.
See ToolCall.
tokenCostnumber | null
This is the cost of the tokens used to produce the actual output.
Example: 0.002
inputTokenCountnumber | null
This is the number of input tokens passed to the LLM model.
Example: 12
outputTokenCountnumber | null
This is the number of output tokens generated by the LLM model.
Example: 3
additionalMetadataRecord<string, unknown> | null
Additional metadata to associate with the golden.
Example: {"source":"faq"}
commentsstring | null
Comments to associate with the golden.
Example: "Reviewed by the support team."
sourceFilestring | null
The source file the golden was retrieved from. Like tags and customColumnKeyValues, this is left unchanged when the request omits it; send null to clear it.
Example: "capitals.csv"
sourceFilesstring[]
The source files the golden was retrieved from. Like tags and customColumnKeyValues, these are left unchanged when the request omits them; send an empty array to clear them.
Example: ["capitals.csv"]
finalizedboolean
Whether the golden is ready to use in evaluations. When pushing or queueing a list of goldens the request decides this for every golden and this field is ignored.
Example: true
customColumnKeyValuesRecord<string, string>
Custom dataset column values keyed by column name. A column that does not exist in the dataset yet is created.
Example: {"difficulty":"easy"}
imagesMappingRecord<string, MLLMImage>
The media this golden refers to, keyed by the id inside each placeholder. Put [DEEPEVAL:IMAGE:<id>] or [DEEPEVAL:PDF:<id>] in a text field where the media belongs, and the platform substitutes the entry with a matching key.
See MLLMImage.
Example: {"map":{"url":"https://example.com/paris.png","local":false}}
tagsstring[]
Tags to associate with the golden, which is useful for grouping and filtering goldens. A tag that does not exist in the dataset yet is created.
Example: ["geography"]
A multi-turn golden to write: the scenario of a conversation with your LLM application and, optionally, its turns.
class MultiTurnGoldenRequest:
scenario: str
expected_outcome: Optional[str] = Field(default=None, alias="expectedOutcome")
user_description: Optional[str] = Field(default=None, alias="userDescription")
turns: Optional[List[Turn]] = None
context: Optional[List[str]] = None
additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
comments: Optional[str] = None
source_file: Optional[str] = Field(default=None, alias="sourceFile")
source_files: Optional[List[str]] = Field(default=None, alias="sourceFiles")
finalized: Optional[bool] = None
custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
tags: Optional[List[str]] = NonescenariostrRequired
This is a description of the conversation context.
Example: "A traveller wants to book a hotel in Paris."
expected_outcomeOptional[str]
This describes the expected outcome, or ideal conversation flow, of the conversation.
Example: "The assistant confirms a reservation near the Louvre."
user_descriptionOptional[str]
This is the description of the user in the conversation.
Example: "A traveller planning a weekend in Paris."
turnsOptional[List[Turn]]
This is the list of turns in the conversation.
See Turn.
Example: [{"role":"user","content":"I need a hotel in Paris near the Louvre."},{"role":"assistant","content":"Hôtel du Louvre has rooms available. Which dates?"}]
contextOptional[List[str]]
This is the context of the conversation.
Example: ["Hôtel du Louvre is a five-minute walk from the museum."]
additional_metadataOptional[Dict[str, Any]]
Additional metadata to associate with the golden.
Example: {"source":"faq"}
commentsOptional[str]
Comments to associate with the golden.
Example: "Reviewed by the support team."
source_fileOptional[str]
The source file the golden was retrieved from. Like tags and customColumnKeyValues, this is left unchanged when the request omits it; send null to clear it.
Example: "capitals.csv"
source_filesOptional[List[str]]
The source files the golden was retrieved from. Like tags and customColumnKeyValues, these are left unchanged when the request omits them; send an empty array to clear them.
Example: ["capitals.csv"]
finalizedOptional[bool]
Whether the golden is ready to use in evaluations. When pushing or queueing a list of goldens the request decides this for every golden and this field is ignored.
Example: true
custom_column_key_valuesOptional[Dict[str, str]]
Custom dataset column values keyed by column name. A column that does not exist in the dataset yet is created.
Example: {"difficulty":"easy"}
images_mappingOptional[Dict[str, MLLMImage]]
The media this golden refers to, keyed by the id inside each placeholder. Put [DEEPEVAL:IMAGE:<id>] or [DEEPEVAL:PDF:<id>] in a text field where the media belongs, and the platform substitutes the entry with a matching key.
See MLLMImage.
Example: {"map":{"url":"https://example.com/paris.png","local":false}}
tagsOptional[List[str]]
Tags to associate with the golden, which is useful for grouping and filtering goldens. A tag that does not exist in the dataset yet is created.
Example: ["geography"]
interface MultiTurnGoldenRequest {
scenario: string;
expectedOutcome?: string | null;
userDescription?: string | null;
turns?: Turn[] | null;
context?: string[] | null;
additionalMetadata?: Record<string, unknown> | null;
comments?: string | null;
sourceFile?: string | null;
sourceFiles?: string[];
finalized?: boolean;
customColumnKeyValues?: Record<string, string>;
imagesMapping?: Record<string, MLLMImage>;
tags?: string[];
}scenariostringRequired
This is a description of the conversation context.
Example: "A traveller wants to book a hotel in Paris."
expectedOutcomestring | null
This describes the expected outcome, or ideal conversation flow, of the conversation.
Example: "The assistant confirms a reservation near the Louvre."
userDescriptionstring | null
This is the description of the user in the conversation.
Example: "A traveller planning a weekend in Paris."
turnsTurn[] | null
This is the list of turns in the conversation.
See Turn.
Example: [{"role":"user","content":"I need a hotel in Paris near the Louvre."},{"role":"assistant","content":"Hôtel du Louvre has rooms available. Which dates?"}]
contextstring[] | null
This is the context of the conversation.
Example: ["Hôtel du Louvre is a five-minute walk from the museum."]
additionalMetadataRecord<string, unknown> | null
Additional metadata to associate with the golden.
Example: {"source":"faq"}
commentsstring | null
Comments to associate with the golden.
Example: "Reviewed by the support team."
sourceFilestring | null
The source file the golden was retrieved from. Like tags and customColumnKeyValues, this is left unchanged when the request omits it; send null to clear it.
Example: "capitals.csv"
sourceFilesstring[]
The source files the golden was retrieved from. Like tags and customColumnKeyValues, these are left unchanged when the request omits them; send an empty array to clear them.
Example: ["capitals.csv"]
finalizedboolean
Whether the golden is ready to use in evaluations. When pushing or queueing a list of goldens the request decides this for every golden and this field is ignored.
Example: true
customColumnKeyValuesRecord<string, string>
Custom dataset column values keyed by column name. A column that does not exist in the dataset yet is created.
Example: {"difficulty":"easy"}
imagesMappingRecord<string, MLLMImage>
The media this golden refers to, keyed by the id inside each placeholder. Put [DEEPEVAL:IMAGE:<id>] or [DEEPEVAL:PDF:<id>] in a text field where the media belongs, and the platform substitutes the entry with a matching key.
See MLLMImage.
Example: {"map":{"url":"https://example.com/paris.png","local":false}}
tagsstring[]
Tags to associate with the golden, which is useful for grouping and filtering goldens. A tag that does not exist in the dataset yet is created.
Example: ["geography"]
MLLMImage
An image referenced from a text field by a [DEEPEVAL:IMAGE:<key>] marker. Send either a public url or the bytes in base64.
class MLLMImage:
url: str
local: bool
base64: Optional[str] = None
filename: Optional[str] = None
mime_type: Optional[str] = Field(default=None, alias="mimeType")
data_base64: Optional[str] = Field(default=None, alias="dataBase64")urlstrRequired
This is the URL of the image.
Example: "https://example.com/everest.png"
localboolRequired
This is true when the image is your local file.
Example: false
base64Optional[str]
The base64 data of the image.
Example: "iVBORw0KGgo="
filenameOptional[str]
The original file name.
Example: "everest.png"
mime_typeOptional[str]
The image's MIME type.
Example: "image/png"
data_base64Optional[str]
The image encoded as a base64 data URL.
Example: "data:image/png;base64,iVBORw0KGgo="
interface MLLMImage {
url: string;
local: boolean;
base64?: string;
filename?: string;
mimeType?: string;
dataBase64?: string;
}urlstringRequired
This is the URL of the image.
Example: "https://example.com/everest.png"
localbooleanRequired
This is true when the image is your local file.
Example: false
base64string
The base64 data of the image.
Example: "iVBORw0KGgo="
filenamestring
The original file name.
Example: "everest.png"
mimeTypestring
The image's MIME type.
Example: "image/png"
dataBase64string
The image encoded as a base64 data URL.
Example: "data:image/png;base64,iVBORw0KGgo="
ToolCall
A tool your LLM application invoked, with what it passed in and what came back.
class ToolCall:
name: str
type: Optional[ToolCallType] = None
description: Optional[str] = None
input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
output: Optional[Any] = None
reasoning: Optional[str] = NonenamestrRequired
This is the name of the tool.
Example: "get_landmark_info"
typeOptional[ToolCallType]
See ToolCallType.
descriptionOptional[str]
This is the description of the tool.
Example: "This tool gives information about a mountain."
input_parametersOptional[Dict[str, Any]]
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputOptional[Any]
This is the output of the tool.
Example: "8,848 metres"
reasoningOptional[str]
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
interface ToolCall {
name: string;
type?: ToolCallType;
description?: string;
inputParameters?: Record<string, unknown> | null;
output?: unknown;
reasoning?: string;
}namestringRequired
This is the name of the tool.
Example: "get_landmark_info"
typeToolCallType
See ToolCallType.
descriptionstring
This is the description of the tool.
Example: "This tool gives information about a mountain."
inputParametersRecord<string, unknown> | null
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputunknown
This is the output of the tool.
Example: "8,848 metres"
reasoningstring
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
ToolCallType
The type of the tool call, either a function or an MCP tool.
class ToolCallType(Enum):
FUNCTION = "FUNCTION"
MCP = "MCP"enum ToolCallType {
FUNCTION = "FUNCTION",
MCP = "MCP",
}FUNCTION · MCP
Turn
One message in a conversation, from either the user or the assistant, with the context and tools behind an assistant reply.
class Turn:
id: Optional[str] = None
role: TurnRole
content: str
user_id: Optional[str] = Field(default=None, alias="userId")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")idOptional[str]
The id of a turn assigned by Confident AI.
Example: "<TURN-ID>"
roleTurnRoleRequired
See TurnRole.
contentstrRequired
The message content of the turn.
Example: "How tall is Mount Everest?"
user_idOptional[str]
The user ID associated with the turn.
Example: "end-user-42"
retrieval_contextOptional[List[str]]
The contexts retrieved to generate the LLM response for this turn.
Example: ["Everest is 8,848 metres tall."]
tools_calledOptional[List[ToolCall]]
The tools called to generate the LLM response for this turn.
See ToolCall.
interface Turn {
id?: string;
role: TurnRole;
content: string;
userId?: string;
retrievalContext?: string[] | null;
toolsCalled?: ToolCall[] | null;
}idstring
The id of a turn assigned by Confident AI.
Example: "<TURN-ID>"
roleTurnRoleRequired
See TurnRole.
contentstringRequired
The message content of the turn.
Example: "How tall is Mount Everest?"
userIdstring
The user ID associated with the turn.
Example: "end-user-42"
retrievalContextstring[] | null
The contexts retrieved to generate the LLM response for this turn.
Example: ["Everest is 8,848 metres tall."]
toolsCalledToolCall[] | null
The tools called to generate the LLM response for this turn.
See ToolCall.
TurnRole
The role of the turn, either user or assistant.
class TurnRole(Enum):
USER = "user"
ASSISTANT = "assistant"enum TurnRole {
USER = "user",
ASSISTANT = "assistant",
}USER · ASSISTANT
Last updated on