Evaluate
Overview
The Confident AI SDK exposes every Evaluate method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
Run Evals
Runs the metrics in metricCollection against your test cases and returns the test run id they were evaluated in. Send either single-turn test cases or multi-turn test cases, not both.
from confident_ai import ConfidentAI
from confident_ai.evaluate import SingleTurnTestCase
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
client = ConfidentAI()
result = client.evaluate.run_evals(
metric_collection="Collection Name",
test_cases=[
SingleTurnTestCase(
input="How tall is mount everest?",
actual_output="No clue, pretty tall I guess?",
expected_output="Mount Everest is 8,848 metres tall.",
retrieval_context=["Everest is 8,848 metres tall."],
tools_called=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
expected_tools=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
context=["Everest is 8,848 metres tall."],
token_cost=0.002,
input_token_count=24,
output_token_count=12,
name="everest-height",
flaky=False,
images_mapping={
"summit": {"url": "https://example.com/everest.png", "local": False}
},
additional_metadata={"region": "Nepal"},
custom_column_key_values={"team": "search"},
tags=["geography"]
)
],
hyperparameters={"model": "gpt-4o-mini"},
identifier="run-399-102",
)For async mode, call a_run_evals and await it as shown below:
result = await client.evaluate.a_run_evals(...)Parameters
| Parameter | Type | Description |
|---|---|---|
metric_collection | str | Required. The name of the metric collection you wish to use for evaluation. |
test_cases | List[TestCase] | Required. This is the list of test cases to evaluate. Every test case in one request must be of the same kind — all single-turn, or all multi-turn. See TestCase. |
hyperparameters | Optional[Dict[str, HyperparameterValue]] | This is any hyperparameters like model or prompt you wish to associate with the test run. See HyperparameterValue. |
identifier | Optional[str] | A unique identifier for the test run. |
import { ConfidentAI } from "confident-ai";
import { ToolCallType } from "confident-ai/common";
const client = new ConfidentAI();
const result = await client.evaluate.runEvals(
"Collection Name",
[
{
input: "How tall is mount everest?",
actualOutput: "No clue, pretty tall I guess?",
expectedOutput: "Mount Everest is 8,848 metres tall.",
retrievalContext: ["Everest is 8,848 metres tall."],
toolsCalled: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
expectedTools: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
context: ["Everest is 8,848 metres tall."],
tokenCost: 0.002,
inputTokenCount: 24,
outputTokenCount: 12,
name: "everest-height",
flaky: false,
imagesMapping: {
summit: { url: "https://example.com/everest.png", local: false }
},
additionalMetadata: { region: "Nepal" },
customColumnKeyValues: { team: "search" },
tags: ["geography"]
}
],
{
hyperparameters: { model: "gpt-4o-mini" },
identifier: "run-399-102"
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
metricCollection | string | Required. The name of the metric collection you wish to use for evaluation. |
testCases | TestCase[] | Required. This is the list of test cases to evaluate. Every test case in one request must be of the same kind — all single-turn, or all multi-turn. See TestCase. |
hyperparameters | Record<string, HyperparameterValue> | This is any hyperparameters like model or prompt you wish to associate with the test run. See HyperparameterValue. |
identifier | string | A unique identifier for the test run. |
Returns
This method returns an object of type EvaluateResult.
Evaluate Span
Queues an evaluation of a span against the metrics in metricCollection. The evaluation runs in the background, and its results are stored on the span, so fetch the span to read them once it has finished.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.evaluate.evaluate_span(
span_uuid="<SPAN-UUID>",
metric_collection="Collection Name",
overwrite_metrics=False,
)For async mode, call a_evaluate_span and await it as shown below:
result = await client.evaluate.a_evaluate_span(...)Parameters
| Parameter | Type | Description |
|---|---|---|
span_uuid | str | Required. The unique identifier of the span. |
metric_collection | str | Required. The name of the single-turn metric collection you wish to use for evaluation. |
overwrite_metrics | Optional[bool] | Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.evaluate.evaluateSpan(
"<SPAN-UUID>",
"Collection Name",
{ overwriteMetrics: false },
);Parameters
| Parameter | Type | Description |
|---|---|---|
spanUuid | string | Required. The unique identifier of the span. |
metricCollection | string | Required. The name of the single-turn metric collection you wish to use for evaluation. |
overwriteMetrics | boolean | Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet. |
Returns
This method returns an object of type EvaluateSpanResult.
Evaluate Thread
Queues an evaluation of a thread against the multi-turn metrics in metricCollection. The evaluation runs in the background, and its results are stored on the thread, so fetch the thread to read them once it has finished.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.evaluate.evaluate_thread(
thread_id="thread-42",
metric_collection="Collection Name",
chatbot_role="A helpful geography assistant.",
overwrite_metrics=False,
)For async mode, call a_evaluate_thread and await it as shown below:
result = await client.evaluate.a_evaluate_thread(...)Parameters
| Parameter | Type | Description |
|---|---|---|
thread_id | str | Required. The id of the thread, as you supplied it when creating its traces. |
metric_collection | str | Required. The name of the multi-turn metric collection you wish to use for evaluation. |
chatbot_role | Optional[str] | This is the role of the chatbot in the thread, which the multi-turn metrics that judge role adherence evaluate the thread against. |
overwrite_metrics | Optional[bool] | Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.evaluate.evaluateThread(
"thread-42",
"Collection Name",
{
chatbotRole: "A helpful geography assistant.",
overwriteMetrics: false
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
threadId | string | Required. The id of the thread, as you supplied it when creating its traces. |
metricCollection | string | Required. The name of the multi-turn metric collection you wish to use for evaluation. |
chatbotRole | string | This is the role of the chatbot in the thread, which the multi-turn metrics that judge role adherence evaluate the thread against. |
overwriteMetrics | boolean | Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet. |
Returns
This method returns an object of type EvaluateThreadResult.
Evaluate Trace
Queues an evaluation of a trace against the metrics in metricCollection. The evaluation runs in the background, and its results are stored on the trace, so fetch the trace to read them once it has finished.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.evaluate.evaluate_trace(
trace_uuid="<TRACE-UUID>",
metric_collection="Collection Name",
overwrite_metrics=False,
)For async mode, call a_evaluate_trace and await it as shown below:
result = await client.evaluate.a_evaluate_trace(...)Parameters
| Parameter | Type | Description |
|---|---|---|
trace_uuid | str | Required. The unique identifier of the trace. |
metric_collection | str | Required. The name of the single-turn metric collection you wish to use for evaluation. |
overwrite_metrics | Optional[bool] | Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.evaluate.evaluateTrace(
"<TRACE-UUID>",
"Collection Name",
{ overwriteMetrics: false },
);Parameters
| Parameter | Type | Description |
|---|---|---|
traceUuid | string | Required. The unique identifier of the trace. |
metricCollection | string | Required. The name of the single-turn metric collection you wish to use for evaluation. |
overwriteMetrics | boolean | Set this to true to re-run every metric in the collection and replace the results already stored, and omit this field to keep those results and only run the metrics that have none yet. |
Returns
This method returns an object of type EvaluateTraceResult.
Types
EvaluateResult
class EvaluateResult:
id: stridstrRequired
This is the unique ID for the test run. This ID is generated by Confident AI and is not to be confused with the identifier provided by the user.
Example: "<TEST-RUN-ID>"
interface EvaluateResult {
id: string;
}idstringRequired
This is the unique ID for the test run. This ID is generated by Confident AI and is not to be confused with the identifier provided by the user.
Example: "<TEST-RUN-ID>"
EvaluateSpanResult
The span whose evaluation was queued.
class EvaluateSpanResult:
id: stridstrRequired
This is the uuid of the span the evaluation was queued for.
Example: "<SPAN-UUID>"
interface EvaluateSpanResult {
id: string;
}idstringRequired
This is the uuid of the span the evaluation was queued for.
Example: "<SPAN-UUID>"
EvaluateThreadResult
The thread whose evaluation was queued.
class EvaluateThreadResult:
id: stridstrRequired
This is the id of the thread the evaluation was queued for.
Example: "thread-42"
interface EvaluateThreadResult {
id: string;
}idstringRequired
This is the id of the thread the evaluation was queued for.
Example: "thread-42"
EvaluateTraceResult
The trace whose evaluation was queued.
class EvaluateTraceResult:
id: stridstrRequired
This is the uuid of the trace the evaluation was queued for.
Example: "<TRACE-UUID>"
interface EvaluateTraceResult {
id: string;
}idstringRequired
This is the uuid of the trace the evaluation was queued for.
Example: "<TRACE-UUID>"
HyperparameterValue
A plain value such as a model name, or a reference to the prompt the run used.
HyperparameterValue = Union[
PromptHyperparameter,
]type HyperparameterValue =
| PromptHyperparameter;A HyperparameterValue is one of the shapes below. Send the fields of one of them, never a mix of both.
A reference to the prompt version the run used.
class PromptHyperparameter:
id: str
type: PromptTypeidstrRequired
This is the id of the prompt version the run used.
Example: "<PROMPT-VERSION-ID>"
typePromptTypeRequired
See PromptType.
interface PromptHyperparameter {
id: string;
type: PromptType;
}idstringRequired
This is the id of the prompt version the run used.
Example: "<PROMPT-VERSION-ID>"
typePromptTypeRequired
See PromptType.
MLLMImage
An image referenced from a text field by a [DEEPEVAL:IMAGE:<key>] marker. Send either a public url or the bytes in base64.
class MLLMImage:
url: str
local: bool
base64: Optional[str] = None
filename: Optional[str] = None
mime_type: Optional[str] = Field(default=None, alias="mimeType")
data_base64: Optional[str] = Field(default=None, alias="dataBase64")urlstrRequired
This is the URL of the image.
Example: "https://example.com/everest.png"
localboolRequired
This is true when the image is your local file.
Example: false
base64Optional[str]
The base64 data of the image.
Example: "iVBORw0KGgo="
filenameOptional[str]
The original file name.
Example: "everest.png"
mime_typeOptional[str]
The image's MIME type.
Example: "image/png"
data_base64Optional[str]
The image encoded as a base64 data URL.
Example: "data:image/png;base64,iVBORw0KGgo="
interface MLLMImage {
url: string;
local: boolean;
base64?: string;
filename?: string;
mimeType?: string;
dataBase64?: string;
}urlstringRequired
This is the URL of the image.
Example: "https://example.com/everest.png"
localbooleanRequired
This is true when the image is your local file.
Example: false
base64string
The base64 data of the image.
Example: "iVBORw0KGgo="
filenamestring
The original file name.
Example: "everest.png"
mimeTypestring
The image's MIME type.
Example: "image/png"
dataBase64string
The image encoded as a base64 data URL.
Example: "data:image/png;base64,iVBORw0KGgo="
PromptType
This is the type of the prompt, which can be either a simple text or a list of messages.
class PromptType(Enum):
TEXT = "TEXT"
LIST = "LIST"enum PromptType {
TEXT = "TEXT",
LIST = "LIST",
}TEXT · LIST
TestCase
One test case to evaluate: single-turn when it carries input, multi-turn when it carries turns. A test case cannot be both, and one request cannot mix the two kinds.
TestCase = Union[
SingleTurnTestCase,
MultiTurnTestCase,
]type TestCase =
| SingleTurnTestCase
| MultiTurnTestCase;A TestCase is one of the shapes below. Send the fields of one of them, never a mix of both.
A test case for a single exchange with your LLM application.
class SingleTurnTestCase:
input: str
actual_output: Optional[str] = Field(default=None, alias="actualOutput")
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
context: Optional[List[str]] = None
token_cost: Optional[float] = Field(default=None, alias="tokenCost")
input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
name: Optional[str] = None
flaky: Optional[bool] = None
images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
tags: Optional[List[str]] = NoneinputstrRequired
This is the input to your LLM application.
Example: "How tall is mount everest?"
actual_outputOptional[str]
This is the actual output of your LLM application.
Example: "No clue, pretty tall I guess?"
expected_outputOptional[str]
This is the expected output of your LLM application, which is the ideal actual output.
Example: "Mount Everest is 8,848 metres tall."
retrieval_contextOptional[List[str]]
This is the retrieval context of your LLM application.
Example: ["Everest is 8,848 metres tall."]
tools_calledOptional[List[ToolCall]]
This is the tools called by your LLM application.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the LLM application.
See ToolCall.
contextOptional[List[str]]
This is the ideal retrieval context of your LLM application.
Example: ["Everest is 8,848 metres tall."]
token_costOptional[float]
This is the cost of the tokens used by the LLM model.
Example: 0.002
input_token_countOptional[int]
This is the number of input tokens passed to the LLM model.
Example: 24
output_token_countOptional[int]
This is the number of output tokens generated by the LLM model.
Example: 12
nameOptional[str]
This is the name of your test case, it allows you to search and match test cases across different test runs.
Example: "everest-height"
flakyOptional[bool]
This is true if the test case's verdict was non-deterministic across runs.
Example: false
images_mappingOptional[Dict[str, MLLMImage]]
This is the mapping of image placeholders in your test case to the images they refer to.
See MLLMImage.
Example: {"summit":{"url":"https://example.com/everest.png","local":false}}
additional_metadataOptional[Dict[str, Any]]
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
custom_column_key_valuesOptional[Dict[str, str]]
This is the custom column key values of the LLM application.
Example: {"team":"search"}
tagsOptional[List[str]]
This is the list of tags associated with the test case, which is useful for grouping and filtering for test cases.
Example: ["geography"]
interface SingleTurnTestCase {
input: string;
actualOutput?: string;
expectedOutput?: string;
retrievalContext?: string[];
toolsCalled?: ToolCall[];
expectedTools?: ToolCall[];
context?: string[];
tokenCost?: number;
inputTokenCount?: number;
outputTokenCount?: number;
name?: string;
flaky?: boolean;
imagesMapping?: Record<string, MLLMImage>;
additionalMetadata?: Record<string, unknown>;
customColumnKeyValues?: Record<string, string>;
tags?: string[];
}inputstringRequired
This is the input to your LLM application.
Example: "How tall is mount everest?"
actualOutputstring
This is the actual output of your LLM application.
Example: "No clue, pretty tall I guess?"
expectedOutputstring
This is the expected output of your LLM application, which is the ideal actual output.
Example: "Mount Everest is 8,848 metres tall."
retrievalContextstring[]
This is the retrieval context of your LLM application.
Example: ["Everest is 8,848 metres tall."]
toolsCalledToolCall[]
This is the tools called by your LLM application.
See ToolCall.
expectedToolsToolCall[]
This is the expected tools to be called by the LLM application.
See ToolCall.
contextstring[]
This is the ideal retrieval context of your LLM application.
Example: ["Everest is 8,848 metres tall."]
tokenCostnumber
This is the cost of the tokens used by the LLM model.
Example: 0.002
inputTokenCountnumber
This is the number of input tokens passed to the LLM model.
Example: 24
outputTokenCountnumber
This is the number of output tokens generated by the LLM model.
Example: 12
namestring
This is the name of your test case, it allows you to search and match test cases across different test runs.
Example: "everest-height"
flakyboolean
This is true if the test case's verdict was non-deterministic across runs.
Example: false
imagesMappingRecord<string, MLLMImage>
This is the mapping of image placeholders in your test case to the images they refer to.
See MLLMImage.
Example: {"summit":{"url":"https://example.com/everest.png","local":false}}
additionalMetadataRecord<string, unknown>
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
customColumnKeyValuesRecord<string, string>
This is the custom column key values of the LLM application.
Example: {"team":"search"}
tagsstring[]
This is the list of tags associated with the test case, which is useful for grouping and filtering for test cases.
Example: ["geography"]
A test case for a conversation with your LLM application.
class MultiTurnTestCase:
turns: List[Turn]
scenario: Optional[str] = None
expected_outcome: Optional[str] = Field(default=None, alias="expectedOutcome")
user_description: Optional[str] = Field(default=None, alias="userDescription")
chatbot_role: Optional[str] = Field(default=None, alias="chatbotRole")
context: Optional[List[str]] = None
token_cost: Optional[float] = Field(default=None, alias="tokenCost")
input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
name: Optional[str] = None
flaky: Optional[bool] = None
images_mapping: Optional[Dict[str, MLLMImage]] = Field(default=None, alias="imagesMapping")
additional_metadata: Optional[Dict[str, Any]] = Field(default=None, alias="additionalMetadata")
custom_column_key_values: Optional[Dict[str, str]] = Field(default=None, alias="customColumnKeyValues")
tags: Optional[List[str]] = NoneturnsList[Turn]Required
This is the list of turns in the conversation.
See Turn.
Example: [{"role":"user","content":"How tall is Mount Everest?"},{"role":"assistant","content":"Mount Everest is 8,848 metres tall."}]
scenarioOptional[str]
This is a description of the conversation context.
Example: "A traveller asking about mountain heights."
expected_outcomeOptional[str]
This describes the expected outcome, or ideal conversation flow, of the conversation.
Example: "The assistant states Everest's height."
user_descriptionOptional[str]
This is the description of the user in the conversation.
Example: "A traveller planning a trek."
chatbot_roleOptional[str]
This is the role of the chatbot in the conversation.
Example: "A helpful geography assistant."
contextOptional[List[str]]
This is the ideal retrieval context of your LLM application.
Example: ["Everest is 8,848 metres tall."]
token_costOptional[float]
This is the cost of the tokens used by the LLM model.
Example: 0.002
input_token_countOptional[int]
This is the number of input tokens passed to the LLM model.
Example: 24
output_token_countOptional[int]
This is the number of output tokens generated by the LLM model.
Example: 12
nameOptional[str]
This is the name of your test case, it allows you to search and match test cases across different test runs.
Example: "everest-height"
flakyOptional[bool]
This is true if the test case's verdict was non-deterministic across runs.
Example: false
images_mappingOptional[Dict[str, MLLMImage]]
This is the mapping of image placeholders in your test case to the images they refer to.
See MLLMImage.
Example: {"summit":{"url":"https://example.com/everest.png","local":false}}
additional_metadataOptional[Dict[str, Any]]
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
custom_column_key_valuesOptional[Dict[str, str]]
This is the custom column key values of the LLM application.
Example: {"team":"search"}
tagsOptional[List[str]]
This is the list of tags associated with the test case, which is useful for grouping and filtering for test cases.
Example: ["geography"]
interface MultiTurnTestCase {
turns: Turn[];
scenario?: string;
expectedOutcome?: string;
userDescription?: string;
chatbotRole?: string;
context?: string[];
tokenCost?: number;
inputTokenCount?: number;
outputTokenCount?: number;
name?: string;
flaky?: boolean;
imagesMapping?: Record<string, MLLMImage>;
additionalMetadata?: Record<string, unknown>;
customColumnKeyValues?: Record<string, string>;
tags?: string[];
}turnsTurn[]Required
This is the list of turns in the conversation.
See Turn.
Example: [{"role":"user","content":"How tall is Mount Everest?"},{"role":"assistant","content":"Mount Everest is 8,848 metres tall."}]
scenariostring
This is a description of the conversation context.
Example: "A traveller asking about mountain heights."
expectedOutcomestring
This describes the expected outcome, or ideal conversation flow, of the conversation.
Example: "The assistant states Everest's height."
userDescriptionstring
This is the description of the user in the conversation.
Example: "A traveller planning a trek."
chatbotRolestring
This is the role of the chatbot in the conversation.
Example: "A helpful geography assistant."
contextstring[]
This is the ideal retrieval context of your LLM application.
Example: ["Everest is 8,848 metres tall."]
tokenCostnumber
This is the cost of the tokens used by the LLM model.
Example: 0.002
inputTokenCountnumber
This is the number of input tokens passed to the LLM model.
Example: 24
outputTokenCountnumber
This is the number of output tokens generated by the LLM model.
Example: 12
namestring
This is the name of your test case, it allows you to search and match test cases across different test runs.
Example: "everest-height"
flakyboolean
This is true if the test case's verdict was non-deterministic across runs.
Example: false
imagesMappingRecord<string, MLLMImage>
This is the mapping of image placeholders in your test case to the images they refer to.
See MLLMImage.
Example: {"summit":{"url":"https://example.com/everest.png","local":false}}
additionalMetadataRecord<string, unknown>
Additional metadata associated with this test case.
Example: {"region":"Nepal"}
customColumnKeyValuesRecord<string, string>
This is the custom column key values of the LLM application.
Example: {"team":"search"}
tagsstring[]
This is the list of tags associated with the test case, which is useful for grouping and filtering for test cases.
Example: ["geography"]
ToolCall
A tool your LLM application invoked, with what it passed in and what came back.
class ToolCall:
name: str
type: Optional[ToolCallType] = None
description: Optional[str] = None
input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
output: Optional[Any] = None
reasoning: Optional[str] = NonenamestrRequired
This is the name of the tool.
Example: "get_landmark_info"
typeOptional[ToolCallType]
See ToolCallType.
descriptionOptional[str]
This is the description of the tool.
Example: "This tool gives information about a mountain."
input_parametersOptional[Dict[str, Any]]
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputOptional[Any]
This is the output of the tool.
Example: "8,848 metres"
reasoningOptional[str]
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
interface ToolCall {
name: string;
type?: ToolCallType;
description?: string;
inputParameters?: Record<string, unknown> | null;
output?: unknown;
reasoning?: string;
}namestringRequired
This is the name of the tool.
Example: "get_landmark_info"
typeToolCallType
See ToolCallType.
descriptionstring
This is the description of the tool.
Example: "This tool gives information about a mountain."
inputParametersRecord<string, unknown> | null
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputunknown
This is the output of the tool.
Example: "8,848 metres"
reasoningstring
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
ToolCallType
The type of the tool call, either a function or an MCP tool.
class ToolCallType(Enum):
FUNCTION = "FUNCTION"
MCP = "MCP"enum ToolCallType {
FUNCTION = "FUNCTION",
MCP = "MCP",
}FUNCTION · MCP
Turn
One message in a conversation, from either the user or the assistant, with the context and tools behind an assistant reply.
class Turn:
id: Optional[str] = None
role: TurnRole
content: str
user_id: Optional[str] = Field(default=None, alias="userId")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")idOptional[str]
The id of a turn assigned by Confident AI.
Example: "<TURN-ID>"
roleTurnRoleRequired
See TurnRole.
contentstrRequired
The message content of the turn.
Example: "How tall is Mount Everest?"
user_idOptional[str]
The user ID associated with the turn.
Example: "end-user-42"
retrieval_contextOptional[List[str]]
The contexts retrieved to generate the LLM response for this turn.
Example: ["Everest is 8,848 metres tall."]
tools_calledOptional[List[ToolCall]]
The tools called to generate the LLM response for this turn.
See ToolCall.
interface Turn {
id?: string;
role: TurnRole;
content: string;
userId?: string;
retrievalContext?: string[] | null;
toolsCalled?: ToolCall[] | null;
}idstring
The id of a turn assigned by Confident AI.
Example: "<TURN-ID>"
roleTurnRoleRequired
See TurnRole.
contentstringRequired
The message content of the turn.
Example: "How tall is Mount Everest?"
userIdstring
The user ID associated with the turn.
Example: "end-user-42"
retrievalContextstring[] | null
The contexts retrieved to generate the LLM response for this turn.
Example: ["Everest is 8,848 metres tall."]
toolsCalledToolCall[] | null
The tools called to generate the LLM response for this turn.
See ToolCall.
TurnRole
The role of the turn, either user or assistant.
class TurnRole(Enum):
USER = "user"
ASSISTANT = "assistant"enum TurnRole {
USER = "user",
ASSISTANT = "assistant",
}USER · ASSISTANT
Last updated on