Traces
Overview
The Confident AI SDK exposes every Trace method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
List Traces
Lists the traces in your Confident AI project one page at a time, newest first by default, as summary rows. Retrieve a trace by uuid for its full detail.
from confident_ai import ConfidentAI
from confident_ai.common import Environment
from confident_ai.traces import TraceSortBy
client = ConfidentAI()
result = client.traces.list(
page_size=25,
cursor="<NEXT-CURSOR>",
start="2025-01-01T00:00:00+00:00",
end="2025-01-31T23:59:59+00:00",
ascending="false",
sort_by=TraceSortBy.CREATEDAT,
environment=Environment.PRODUCTION,
metadata={"client": "acme-corp"},
)For async mode, call a_list and await it as shown below:
result = await client.traces.a_list(...)Parameters
| Parameter | Type | Description |
|---|---|---|
page_size | Optional[int] | The number of results per page, at most 100. Defaults to 25. |
cursor | Optional[str] | This is used for pagination, and should be set to the nextCursor value returned in the previous response to get the next page of results. |
start | Optional[str] | This filters for results created at or after the specified start datetime, in ISO 8601 format. Defaults to 60 days ago. |
end | Optional[str] | This filters for results created before the specified end datetime, in ISO 8601 format. Defaults to the current time. |
ascending | Optional[Literal['true', 'false']] | This determines if the field specified in sortBy should be in ascending order. Defaults to false, which returns the newest results first. |
sort_by | Optional[TraceSortBy] | This determines the field to sort by. Defaults to createdAt. See TraceSortBy. |
environment | Optional[Environment] | This filters the traces by the environment where the trace was created, and returns traces from all environments if not specified. See Environment. |
metadata | Optional[Dict[str, str]] | Filter traces by metadata key-value pairs using bracket notation, for example metadata[client]=acme-corp. Every pair must match. |
import { ConfidentAI } from "confident-ai";
import { Environment } from "confident-ai/common";
import { TraceSortBy } from "confident-ai/traces";
const client = new ConfidentAI();
const result = await client.traces.list(
{
pageSize: 25,
cursor: "<NEXT-CURSOR>",
start: "2025-01-01T00:00:00+00:00",
end: "2025-01-31T23:59:59+00:00",
ascending: "false",
sortBy: TraceSortBy.CREATEDAT,
environment: Environment.PRODUCTION,
metadata: { client: "acme-corp" }
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
pageSize | number | The number of results per page, at most 100. Defaults to 25. |
cursor | string | This is used for pagination, and should be set to the nextCursor value returned in the previous response to get the next page of results. |
start | string | This filters for results created at or after the specified start datetime, in ISO 8601 format. Defaults to 60 days ago. |
end | string | This filters for results created before the specified end datetime, in ISO 8601 format. Defaults to the current time. |
ascending | "true" | "false" | This determines if the field specified in sortBy should be in ascending order. Defaults to false, which returns the newest results first. |
sortBy | TraceSortBy | This determines the field to sort by. Defaults to createdAt. See TraceSortBy. |
environment | Environment | This filters the traces by the environment where the trace was created, and returns traces from all environments if not specified. See Environment. |
metadata | Record<string, string> | Filter traces by metadata key-value pairs using bracket notation, for example metadata[client]=acme-corp. Every pair must match. |
Returns
This method returns an object of type TraceList.
Create Trace
Creates a trace in your Confident AI project, along with the spans it contains, and returns its uuid. A uuid that is not a UUID is hashed into one, so the returned value is what every later lookup uses.
from confident_ai import ConfidentAI
from confident_ai.traces import CustomerRequest
from confident_ai.common import Environment
from confident_ai.common import EvaluationErrorType
from confident_ai.traces import MetricDataConfig
from confident_ai.traces import ThreadRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
from confident_ai.common import TraceSpanStatus
from confident_ai.traces import UserRequest
client = ConfidentAI()
result = client.traces.create(
uuid="<TRACE-UUID>",
start_time="2025-01-15T10:30:00+00:00",
end_time="2025-01-15T10:30:05+00:00",
name="Geography QA",
input="What is the capital of France?",
output="The capital of France is Paris.",
status=TraceSpanStatus.SUCCESS,
environment=Environment.PRODUCTION,
metadata={"client": "acme-corp"},
tags=["geography"],
thread_id="thread-42",
thread=ThreadRequest(
id="thread-42",
metadata={"client": "acme-corp", "agentId": "geography-agent"},
tags=["vip"]
),
user_id="end-user-42",
user=UserRequest(
id="end-user-42",
name="Marta Ruiz"
),
customer_id="acme-hotels",
customer=CustomerRequest(
id="acme-hotels",
name="Acme Hotels"
),
metric_collection="Collection Name",
test_run_id="<TEST-RUN-ID>",
test_case_id="<TEST-CASE-ID>",
turn_id="<TURN-ID>",
retrieval_context=[
"Paris is the capital and most populous city of France."
],
context=["Paris is the capital of France."],
expected_output="Paris",
tools_called=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
expected_tools=[
ToolCall(
name="get_landmark_info",
type=ToolCallType.FUNCTION,
description="This tool gives information about a mountain.",
input_parameters={"mountain": "Everest"},
output="8,848 metres",
reasoning="The user asked for the height of a mountain."
)
],
spans=[
{
"uuid": "<SPAN-UUID>",
"type": "LLM",
"name": "OpenAI Call",
"model": "gpt-4o",
"provider": "OpenAI",
"integration": "LangChain",
"input": "What is the capital of France?",
"output": "The capital of France is Paris.",
"startTime": "2025-01-15T10:30:00+00:00",
"endTime": "2025-01-15T10:30:02+00:00"
}
],
metrics_data=[
MetricDataConfig(
name="Answer Relevancy",
score=0.95,
success=True,
threshold=0.5,
strict_mode=False,
flaky=False,
reason="The answer directly states the capital of France.",
evaluation_model="gpt-4o",
evaluation_cost=0.0004,
error="<ERROR>",
error_type=EvaluationErrorType.AI_CONNECTION_ERROR,
verbose_logs="<VERBOSE-LOGS>"
)
],
attachments={
"doc-1": {"mimeType": "application/pdf", "dataBase64": "JVBERi0xLjQK"}
},
)For async mode, call a_create and await it as shown below:
result = await client.traces.a_create(...)Parameters
| Parameter | Type | Description |
|---|---|---|
uuid | str | Required. The unique identifier of the trace, generated by your application. Values that are not UUIDs are hashed into one, and the hashed uuid is what the response and every later lookup use. |
start_time | str | Required. This is the time the trace started, as an ISO 8601 datetime. |
end_time | str | Required. This is the time the trace ended, as an ISO 8601 datetime. |
name | Optional[str] | This is the name of the trace. |
input | Optional[Any] | This is the input to the trace, as a string or any JSON value. |
output | Optional[Any] | This is the output of the trace, as a string or any JSON value. |
status | Optional[TraceSpanStatus] | See TraceSpanStatus. |
environment | Optional[Environment] | See Environment. |
metadata | Optional[Dict[str, Any]] | This is any additional metadata associated with the trace. |
tags | Optional[List[str]] | This is any tags associated with the trace, which helps with grouping traces and filtering them on the Confident AI platform. |
thread_id | Optional[str] | This is the unique identifier of the thread associated with the trace, which groups traces in the same thread into a conversation. |
thread | Optional[ThreadRequest] | See ThreadRequest. |
user_id | Optional[str] | This is the unique identifier for your end user for the trace. |
user | Optional[UserRequest] | See UserRequest. |
customer_id | Optional[str] | This is the unique identifier of the customer the trace belongs to — the account, tenant or organization your end user belongs to. |
customer | Optional[CustomerRequest] | See CustomerRequest. |
metric_collection | Optional[str] | This is the metric collection you wish to use to evaluate the trace. |
test_run_id | Optional[str] | This is the unique identifier of the test run to associate the trace with. When set, the trace becomes one test case in that test run and metricCollection is required. It cannot be combined with testCaseId. |
test_case_id | Optional[str] | The id of an existing test case to attach the trace to, when the trace was produced while evaluating that test case. |
turn_id | Optional[str] | The id of the conversational test case turn to attach the trace to, when the trace was produced while evaluating that turn. |
retrieval_context | Optional[List[str]] | This is the retrieval context of your trace, which is to be used for evaluation. |
context | Optional[List[str]] | This is the ideal retrieval context of your trace, which is to be used for evaluation. |
expected_output | Optional[str] | This is the expected output of your trace, which is the ideal actual output and to be used for evaluation. |
tools_called | Optional[List[ToolCall]] | This is the tools called by your trace, which is to be used for evaluation. See ToolCall. |
expected_tools | Optional[List[ToolCall]] | This is the expected tools to be called by the trace, which is to be used for evaluation. See ToolCall. |
spans | Optional[List[SpanRequest]] | This is the list of spans in the trace. Each span's type decides which fields it accepts. See SpanRequest. |
metrics_data | Optional[List[MetricDataConfig]] | Metric results you already computed for this trace, recorded as-is instead of being evaluated by Confident AI. See MetricDataConfig. |
attachments | Optional[Dict[str, TraceAttachment]] | Map of attachment ids to payloads for all [DEEPEVAL:IMAGE:…] and [DEEPEVAL:PDF:…] markers in this trace. Define attachments at the trace level with the same ids for the same instances. See TraceAttachment. |
import { ConfidentAI } from "confident-ai";
import {
Environment,
EvaluationErrorType,
ToolCallType,
TraceSpanStatus,
} from "confident-ai/common";
const client = new ConfidentAI();
const result = await client.traces.create(
"<TRACE-UUID>",
"2025-01-15T10:30:00+00:00",
"2025-01-15T10:30:05+00:00",
{
name: "Geography QA",
input: "What is the capital of France?",
output: "The capital of France is Paris.",
status: TraceSpanStatus.SUCCESS,
environment: Environment.PRODUCTION,
metadata: { client: "acme-corp" },
tags: ["geography"],
threadId: "thread-42",
thread: {
id: "thread-42",
metadata: { client: "acme-corp", agentId: "geography-agent" },
tags: ["vip"]
},
userId: "end-user-42",
user: {
id: "end-user-42",
name: "Marta Ruiz"
},
customerId: "acme-hotels",
customer: {
id: "acme-hotels",
name: "Acme Hotels"
},
metricCollection: "Collection Name",
testRunId: "<TEST-RUN-ID>",
testCaseId: "<TEST-CASE-ID>",
turnId: "<TURN-ID>",
retrievalContext: [
"Paris is the capital and most populous city of France."
],
context: ["Paris is the capital of France."],
expectedOutput: "Paris",
toolsCalled: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
expectedTools: [
{
name: "get_landmark_info",
type: ToolCallType.FUNCTION,
description: "This tool gives information about a mountain.",
inputParameters: { mountain: "Everest" },
output: "8,848 metres",
reasoning: "The user asked for the height of a mountain."
}
],
spans: [
{
uuid: "<SPAN-UUID>",
type: "LLM",
name: "OpenAI Call",
model: "gpt-4o",
provider: "OpenAI",
integration: "LangChain",
input: "What is the capital of France?",
output: "The capital of France is Paris.",
startTime: "2025-01-15T10:30:00+00:00",
endTime: "2025-01-15T10:30:02+00:00"
}
],
metricsData: [
{
name: "Answer Relevancy",
score: 0.95,
success: true,
threshold: 0.5,
strictMode: false,
flaky: false,
reason: "The answer directly states the capital of France.",
evaluationModel: "gpt-4o",
evaluationCost: 0.0004,
error: "<ERROR>",
errorType: EvaluationErrorType.AI_CONNECTION_ERROR,
verboseLogs: "<VERBOSE-LOGS>"
}
],
attachments: {
doc-1: { mimeType: "application/pdf", dataBase64: "JVBERi0xLjQK" }
}
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
uuid | string | Required. The unique identifier of the trace, generated by your application. Values that are not UUIDs are hashed into one, and the hashed uuid is what the response and every later lookup use. |
startTime | string | Required. This is the time the trace started, as an ISO 8601 datetime. |
endTime | string | Required. This is the time the trace ended, as an ISO 8601 datetime. |
name | string | This is the name of the trace. |
input | unknown | This is the input to the trace, as a string or any JSON value. |
output | unknown | This is the output of the trace, as a string or any JSON value. |
status | TraceSpanStatus | See TraceSpanStatus. |
environment | Environment | See Environment. |
metadata | Record<string, unknown> | This is any additional metadata associated with the trace. |
tags | string[] | This is any tags associated with the trace, which helps with grouping traces and filtering them on the Confident AI platform. |
threadId | string | This is the unique identifier of the thread associated with the trace, which groups traces in the same thread into a conversation. |
thread | ThreadRequest | See ThreadRequest. |
userId | string | This is the unique identifier for your end user for the trace. |
user | UserRequest | See UserRequest. |
customerId | string | This is the unique identifier of the customer the trace belongs to — the account, tenant or organization your end user belongs to. |
customer | CustomerRequest | See CustomerRequest. |
metricCollection | string | This is the metric collection you wish to use to evaluate the trace. |
testRunId | string | This is the unique identifier of the test run to associate the trace with. When set, the trace becomes one test case in that test run and metricCollection is required. It cannot be combined with testCaseId. |
testCaseId | string | The id of an existing test case to attach the trace to, when the trace was produced while evaluating that test case. |
turnId | string | The id of the conversational test case turn to attach the trace to, when the trace was produced while evaluating that turn. |
retrievalContext | string[] | This is the retrieval context of your trace, which is to be used for evaluation. |
context | string[] | This is the ideal retrieval context of your trace, which is to be used for evaluation. |
expectedOutput | string | This is the expected output of your trace, which is the ideal actual output and to be used for evaluation. |
toolsCalled | ToolCall[] | This is the tools called by your trace, which is to be used for evaluation. See ToolCall. |
expectedTools | ToolCall[] | This is the expected tools to be called by the trace, which is to be used for evaluation. See ToolCall. |
spans | SpanRequest[] | This is the list of spans in the trace. Each span's type decides which fields it accepts. See SpanRequest. |
metricsData | MetricDataConfig[] | Metric results you already computed for this trace, recorded as-is instead of being evaluated by Confident AI. See MetricDataConfig. |
attachments | Record<string, TraceAttachment> | Map of attachment ids to payloads for all [DEEPEVAL:IMAGE:…] and [DEEPEVAL:PDF:…] markers in this trace. Define attachments at the trace level with the same ids for the same instances. See TraceAttachment. |
Returns
This method returns an object of type TraceRef.
Get Trace
Retrieves a trace by uuid from your Confident AI project, with its spans, full input and output, evaluation fields, classifier labels, results and annotations.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.traces.get(trace_uuid="<TRACE-UUID>")For async mode, call a_get and await it as shown below:
result = await client.traces.a_get(...)Parameters
| Parameter | Type | Description |
|---|---|---|
trace_uuid | str | Required. The unique identifier of the trace. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.traces.get("<TRACE-UUID>");Parameters
| Parameter | Type | Description |
|---|---|---|
traceUuid | string | Required. The unique identifier of the trace. |
Returns
This method returns an object of type Trace.
Types
AnnotationFieldType
The kind of value an annotation holds: TEXT, NUMBER, FLOAT or BOOLEAN for a free value, CHOICE or MULTIPLE_CHOICE for a choice from the options in the field's config, and THUMBS_RATING or FIVE_STAR_RATING for a rating.
class AnnotationFieldType(Enum):
TEXT = "TEXT"
NUMBER = "NUMBER"
FLOAT = "FLOAT"
BOOLEAN = "BOOLEAN"
CHOICE = "CHOICE"
MULTIPLE_CHOICE = "MULTIPLE_CHOICE"
FIVE_STAR_RATING = "FIVE_STAR_RATING"
THUMBS_RATING = "THUMBS_RATING"enum AnnotationFieldType {
TEXT = "TEXT",
NUMBER = "NUMBER",
FLOAT = "FLOAT",
BOOLEAN = "BOOLEAN",
CHOICE = "CHOICE",
MULTIPLE_CHOICE = "MULTIPLE_CHOICE",
FIVE_STAR_RATING = "FIVE_STAR_RATING",
THUMBS_RATING = "THUMBS_RATING",
}TEXT · NUMBER · FLOAT · BOOLEAN · CHOICE · MULTIPLE_CHOICE · FIVE_STAR_RATING · THUMBS_RATING
AnnotationSummary
A human rating as it appears under the trace, span or thread it was left on, without repeating the ids of that target.
class AnnotationSummary:
id: str
field_type: AnnotationFieldType = Field(alias="fieldType")
value: Optional[Union[str, float, bool, List[str]]]
name: Optional[str]
explanation: Optional[str]
expected_outcome: Optional[str] = Field(alias="expectedOutcome")
expected_output: Optional[str] = Field(alias="expectedOutput")
created_at: str = Field(alias="createdAt")
user: Optional[UserReference]idstrRequired
This is the id of the annotation generated by Confident AI.
Example: "<ANNOTATION-ID>"
field_typeAnnotationFieldTypeRequired
See AnnotationFieldType.
valueOptional[Union[str, float, bool, List[str]]]Required
The annotated value, in the shape its fieldType expects: a boolean for THUMBS_RATING (true is thumbs up) and BOOLEAN, a whole number from 1 to 5 for FIVE_STAR_RATING, a number for NUMBER or FLOAT, a string for TEXT and CHOICE, and a list of strings for MULTIPLE_CHOICE. Null when the annotation was left without a value.
Example: true
nameOptional[str]Required
The name of the annotation.
explanationOptional[str]Required
This is the explanation for the annotation.
Example: "Correct and concise."
expected_outcomeOptional[str]Required
This is the annotated expected outcome, for conversation annotations.
expected_outputOptional[str]Required
This is the annotated expected output, for span and trace annotations.
Example: "The capital of France is Paris."
created_atstrRequired
The timestamp when the annotation was created.
Example: "2025-01-15T11:00:00+00:00"
userOptional[UserReference]Required
See UserReference.
interface AnnotationSummary {
id: string;
fieldType: AnnotationFieldType;
value: string | number | boolean | string[] | null;
name: string | null;
explanation: string | null;
expectedOutcome: string | null;
expectedOutput: string | null;
createdAt: string;
user: UserReference | null;
}idstringRequired
This is the id of the annotation generated by Confident AI.
Example: "<ANNOTATION-ID>"
fieldTypeAnnotationFieldTypeRequired
See AnnotationFieldType.
valuestring | number | boolean | string[] | nullRequired
The annotated value, in the shape its fieldType expects: a boolean for THUMBS_RATING (true is thumbs up) and BOOLEAN, a whole number from 1 to 5 for FIVE_STAR_RATING, a number for NUMBER or FLOAT, a string for TEXT and CHOICE, and a list of strings for MULTIPLE_CHOICE. Null when the annotation was left without a value.
Example: true
namestring | nullRequired
The name of the annotation.
explanationstring | nullRequired
This is the explanation for the annotation.
Example: "Correct and concise."
expectedOutcomestring | nullRequired
This is the annotated expected outcome, for conversation annotations.
expectedOutputstring | nullRequired
This is the annotated expected output, for span and trace annotations.
Example: "The capital of France is Paris."
createdAtstringRequired
The timestamp when the annotation was created.
Example: "2025-01-15T11:00:00+00:00"
userUserReference | nullRequired
See UserReference.
Classification
A label assigned by one of your project's classifiers, with the reason it was chosen.
class Classification:
label: str
reason: strlabelstrRequired
The label the classifier assigned.
Example: "geography"
reasonstrRequired
The classifier's reason for choosing the label.
Example: "The user asks for the capital city of a country."
interface Classification {
label: string;
reason: string;
}labelstringRequired
The label the classifier assigned.
Example: "geography"
reasonstringRequired
The classifier's reason for choosing the label.
Example: "The user asks for the capital city of a country."
CustomerRequest
Customer-level fields applied to the customer record. id is an alternate way to specify the customer and must match top-level customerId if both are provided. name only takes effect when a customer id is resolvable; successive ingestions take the latest name.
class CustomerRequest:
id: Optional[str] = None
name: Optional[str] = NoneidOptional[str]
The customer id. Equivalent to top-level customerId; if both are set they must match.
Example: "acme-hotels"
nameOptional[str]
A human-readable display name for the customer, shown instead of the id.
Example: "Acme Hotels"
interface CustomerRequest {
id?: string;
name?: string | null;
}idstring
The customer id. Equivalent to top-level customerId; if both are set they must match.
Example: "acme-hotels"
namestring | null
A human-readable display name for the customer, shown instead of the id.
Example: "Acme Hotels"
Environment
This is the environment where your trace was posted, which helps with separating and debugging traces from different environments on the Confident AI platform.
class Environment(Enum):
PRODUCTION = "production"
DEVELOPMENT = "development"
STAGING = "staging"
TESTING = "testing"enum Environment {
PRODUCTION = "production",
DEVELOPMENT = "development",
STAGING = "staging",
TESTING = "testing",
}PRODUCTION · DEVELOPMENT · STAGING · TESTING
EvaluationErrorType
Why an evaluation errored: the AI connection or a transformer failed, the evaluation model failed, the test case lacked the parameters the metric needs, or an internal error occurred.
class EvaluationErrorType(Enum):
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR"
TRANSFORMER_ERROR = "TRANSFORMER_ERROR"
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR"
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS"
INTERNAL_ERROR = "INTERNAL_ERROR"enum EvaluationErrorType {
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR",
TRANSFORMER_ERROR = "TRANSFORMER_ERROR",
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR",
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS",
INTERNAL_ERROR = "INTERNAL_ERROR",
}AI_CONNECTION_ERROR · TRANSFORMER_ERROR · EVALUATION_MODEL_ERROR · INVALID_TEST_CASE_PARAMETERS · INTERNAL_ERROR
MetricData
The result of an evaluated metric, with the ids of whatever it was recorded against.
class MetricData:
id: str
name: str
score: Optional[float]
reason: Optional[str]
success: Optional[bool]
threshold: Optional[float]
strict_mode: bool = Field(alias="strictMode")
skipped: bool
flaky: bool
evaluation_model: Optional[str] = Field(alias="evaluationModel")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
error: Optional[str]
error_type: Optional[EvaluationErrorType] = Field(alias="errorType")
created_at: str = Field(alias="createdAt")
evaluated_at: Optional[str] = Field(alias="evaluatedAt")
multi_turn: bool = Field(alias="multiTurn")
trace_uuid: Optional[str] = Field(alias="traceUuid")
span_uuid: Optional[str] = Field(alias="spanUuid")
thread_id: Optional[str] = Field(alias="threadId")
test_case_id: Optional[str] = Field(alias="testCaseId")
test_run_id: Optional[str] = Field(alias="testRunId")idstrRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestrRequired
The name of the metric.
Example: "Answer Relevancy"
scoreOptional[float]Required
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonOptional[str]Required
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successOptional[bool]Required
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdOptional[float]Required
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strict_modeboolRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedboolRequired
Whether the metric evaluation was skipped.
Example: false
flakyboolRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluation_modelOptional[str]Required
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluation_costOptional[float]Required
The cost of running the evaluation in USD.
Example: 0.0004
errorOptional[str]Required
The error message if the evaluation failed.
error_typeOptional[EvaluationErrorType]Required
See EvaluationErrorType.
created_atstrRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluated_atOptional[str]Required
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
multi_turnboolRequired
Whether this metric was evaluated on a multi-turn conversation.
Example: false
trace_uuidOptional[str]Required
The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
span_uuidOptional[str]Required
The uuid of the span this metric was evaluated on, for component-level metrics.
thread_idOptional[str]Required
The id of the thread this metric was evaluated on, for conversation-level metrics.
test_case_idOptional[str]Required
The id of the test case this metric was evaluated on.
test_run_idOptional[str]Required
The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.
interface MetricData {
id: string;
name: string;
score: number | null;
reason: string | null;
success: boolean | null;
threshold: number | null;
strictMode: boolean;
skipped: boolean;
flaky: boolean;
evaluationModel: string | null;
evaluationCost: number | null;
error: string | null;
errorType: EvaluationErrorType | null;
createdAt: string;
evaluatedAt: string | null;
multiTurn: boolean;
traceUuid: string | null;
spanUuid: string | null;
threadId: string | null;
testCaseId: string | null;
testRunId: string | null;
}idstringRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestringRequired
The name of the metric.
Example: "Answer Relevancy"
scorenumber | nullRequired
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonstring | nullRequired
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successboolean | nullRequired
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdnumber | nullRequired
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strictModebooleanRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedbooleanRequired
Whether the metric evaluation was skipped.
Example: false
flakybooleanRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluationModelstring | nullRequired
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluationCostnumber | nullRequired
The cost of running the evaluation in USD.
Example: 0.0004
errorstring | nullRequired
The error message if the evaluation failed.
errorTypeEvaluationErrorType | nullRequired
See EvaluationErrorType.
createdAtstringRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluatedAtstring | nullRequired
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
multiTurnbooleanRequired
Whether this metric was evaluated on a multi-turn conversation.
Example: false
traceUuidstring | nullRequired
The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
spanUuidstring | nullRequired
The uuid of the span this metric was evaluated on, for component-level metrics.
threadIdstring | nullRequired
The id of the thread this metric was evaluated on, for conversation-level metrics.
testCaseIdstring | nullRequired
The id of the test case this metric was evaluated on.
testRunIdstring | nullRequired
The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.
MetricDataConfig
A metric result you computed yourself, recorded on the trace or span as-is instead of being evaluated by Confident AI.
class MetricDataConfig:
name: str
score: Optional[float] = None
success: Optional[bool] = None
threshold: Optional[float] = None
strict_mode: Optional[bool] = Field(default=None, alias="strictMode")
flaky: Optional[bool] = None
reason: Optional[str] = None
evaluation_model: Optional[str] = Field(default=None, alias="evaluationModel")
evaluation_cost: Optional[float] = Field(default=None, alias="evaluationCost")
error: Optional[str] = None
error_type: Optional[EvaluationErrorType] = Field(default=None, alias="errorType")
verbose_logs: Optional[str] = Field(default=None, alias="verboseLogs")namestrRequired
The name of the metric.
Example: "Answer Relevancy"
scoreOptional[float]
The metric score, typically between 0 and 1.
Example: 0.95
successOptional[bool]
Whether the metric passed its threshold.
Example: true
thresholdOptional[float]
The threshold the metric was scored against.
Example: 0.5
strict_modeOptional[bool]
Whether the metric ran in strict mode, which outputs a binary score of 0 or 1.
Example: false
flakyOptional[bool]
Whether the metric's verdict was non-deterministic across runs.
Example: false
reasonOptional[str]
The reason for the metric score.
Example: "The answer directly states the capital of France."
evaluation_modelOptional[str]
The model used to evaluate the metric.
Example: "gpt-4o"
evaluation_costOptional[float]
The cost of running the evaluation in USD.
Example: 0.0004
errorOptional[str]
The error message if the evaluation failed.
error_typeOptional[EvaluationErrorType]
See EvaluationErrorType.
verbose_logsOptional[str]
Detailed logs from the evaluation.
interface MetricDataConfig {
name: string;
score?: number | null;
success?: boolean | null;
threshold?: number | null;
strictMode?: boolean;
flaky?: boolean;
reason?: string | null;
evaluationModel?: string | null;
evaluationCost?: number | null;
error?: string | null;
errorType?: EvaluationErrorType | null;
verboseLogs?: string | null;
}namestringRequired
The name of the metric.
Example: "Answer Relevancy"
scorenumber | null
The metric score, typically between 0 and 1.
Example: 0.95
successboolean | null
Whether the metric passed its threshold.
Example: true
thresholdnumber | null
The threshold the metric was scored against.
Example: 0.5
strictModeboolean
Whether the metric ran in strict mode, which outputs a binary score of 0 or 1.
Example: false
flakyboolean
Whether the metric's verdict was non-deterministic across runs.
Example: false
reasonstring | null
The reason for the metric score.
Example: "The answer directly states the capital of France."
evaluationModelstring | null
The model used to evaluate the metric.
Example: "gpt-4o"
evaluationCostnumber | null
The cost of running the evaluation in USD.
Example: 0.0004
errorstring | null
The error message if the evaluation failed.
errorTypeEvaluationErrorType | null
See EvaluationErrorType.
verboseLogsstring | null
Detailed logs from the evaluation.
Span
A span with its full input and output, evaluation fields, results and annotations.
class Span:
uuid: str
trace_uuid: str = Field(alias="traceUuid")
parent_uuid: Optional[str] = Field(alias="parentUuid")
name: Optional[str]
type: SpanType
status: TraceSpanStatus
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
error: Optional[str]
integration: Optional[str]
provider: Optional[str]
model: Optional[str]
endpoint: Optional[str]
cost: Optional[float]
input_token_cost: Optional[float] = Field(alias="inputTokenCost")
output_token_cost: Optional[float] = Field(alias="outputTokenCost")
cost_per_input_token: Optional[float] = Field(alias="costPerInputToken")
cost_per_output_token: Optional[float] = Field(alias="costPerOutputToken")
input_token_count: Optional[int] = Field(alias="inputTokenCount")
output_token_count: Optional[int] = Field(alias="outputTokenCount")
prompt_alias: Optional[str] = Field(alias="promptAlias")
prompt_version: Optional[str] = Field(alias="promptVersion")
prompt_label: Optional[str] = Field(alias="promptLabel")
prompt_commit_hash: Optional[str] = Field(alias="promptCommitHash")
embedder: Optional[str]
top_k: Optional[int] = Field(alias="topK")
chunk_size: Optional[int] = Field(alias="chunkSize")
description: Optional[str]
agent_handoffs: Optional[List[str]] = Field(alias="agentHandoffs")
available_tools: Optional[List[str]] = Field(alias="availableTools")
metadata: Optional[Dict[str, Any]]
metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
input: Optional[str]
output: Optional[str]
expected_output: Optional[str] = Field(alias="expectedOutput")
retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
context: Optional[List[str]]
tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
metrics_data: List[MetricData] = Field(alias="metricsData")
annotations: List[AnnotationSummary]uuidstrRequired
This is the unique identifier of the span.
Example: "<SPAN-UUID>"
trace_uuidstrRequired
This is the uuid of the trace containing the span.
Example: "<TRACE-UUID>"
parent_uuidOptional[str]Required
This is the uuid of the parent span, or null for a root span.
Example: "<PARENT-SPAN-UUID>"
nameOptional[str]Required
This is the name of the span.
Example: "OpenAI Call"
typeSpanTypeRequired
See SpanType.
statusTraceSpanStatusRequired
See TraceSpanStatus.
start_timestrRequired
This is the time the span started.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the span ended.
Example: "2025-01-15T10:30:02+00:00"
errorOptional[str]Required
This is the error string that caused the span to fail, or null when no error occurred.
integrationOptional[str]Required
This is the integration associated with the span.
Example: "LangChain"
providerOptional[str]Required
This is the LLM provider used in an LLM span.
Example: "OpenAI"
modelOptional[str]Required
This is the LLM model used in an LLM span.
Example: "gpt-4o"
endpointOptional[str]Required
This is the API endpoint the model was called through in an LLM span.
costOptional[float]Required
This is the total cost of the span in USD, or null when it is not known.
Example: 0.00018
input_token_costOptional[float]Required
This is the total cost of the input tokens passed to the LLM model in an LLM span.
Example: 0.00006
output_token_costOptional[float]Required
This is the total cost of the output tokens generated by the LLM model in an LLM span.
Example: 0.00012
cost_per_input_tokenOptional[float]Required
This is the cost per input token of the LLM model for an LLM span.
Example: 0.0000025
cost_per_output_tokenOptional[float]Required
This is the cost per output token of the LLM model for an LLM span.
Example: 0.00001
input_token_countOptional[int]Required
This is the total number of input tokens passed to the LLM model in an LLM span.
Example: 24
output_token_countOptional[int]Required
This is the total number of output tokens generated by the LLM model in an LLM span.
Example: 12
prompt_aliasOptional[str]Required
This is the alias of your prompt which is stored on Confident AI.
Example: "geography-assistant"
prompt_versionOptional[str]Required
This is the version assigned to your prompt on Confident AI.
Example: "00.00.01"
prompt_labelOptional[str]Required
This is the label assigned to a specific version of prompt on the Confident AI platform.
Example: "production"
prompt_commit_hashOptional[str]Required
This is the hash of the current prompt being logged in the llm span.
Example: "bab04ce"
embedderOptional[str]Required
This is the embedder model used in a retriever span.
top_kOptional[int]Required
This is the top K chunks retrieved from your knowledge base in a retriever span.
chunk_sizeOptional[int]Required
This is the chunk size of each retrieved context for a retriever span.
descriptionOptional[str]Required
This is a description if the span is a tool span.
agent_handoffsOptional[List[str]]Required
This is the list of agent handoffs associated with an agent span.
available_toolsOptional[List[str]]Required
This is the list of available tools associated with an agent span.
metadataOptional[Dict[str, Any]]Required
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metric_collection_nameOptional[str]Required
This is the name of the metric collection to evaluate the span.
Example: "LLM Collection Name"
inputOptional[str]Required
This is the input to the span. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputOptional[str]Required
This is the output of the span. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expected_outputOptional[str]Required
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
retrieval_contextOptional[List[str]]Required
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]Required
This is the ideal retrieval context of your span, which is to be used for evaluation.
tools_calledOptional[List[ToolCall]]Required
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]Required
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
metrics_dataList[MetricData]Required
This is the metrics data associated with the span.
See MetricData.
annotationsList[AnnotationSummary]Required
This is the list of annotations associated with the span.
See AnnotationSummary.
interface Span {
uuid: string;
traceUuid: string;
parentUuid: string | null;
name: string | null;
type: SpanType;
status: TraceSpanStatus;
startTime: string;
endTime: string;
error: string | null;
integration: string | null;
provider: string | null;
model: string | null;
endpoint: string | null;
cost: number | null;
inputTokenCost: number | null;
outputTokenCost: number | null;
costPerInputToken: number | null;
costPerOutputToken: number | null;
inputTokenCount: number | null;
outputTokenCount: number | null;
promptAlias: string | null;
promptVersion: string | null;
promptLabel: string | null;
promptCommitHash: string | null;
embedder: string | null;
topK: number | null;
chunkSize: number | null;
description: string | null;
agentHandoffs: string[] | null;
availableTools: string[] | null;
metadata: Record<string, unknown> | null;
metricCollectionName: string | null;
input: string | null;
output: string | null;
expectedOutput: string | null;
retrievalContext: string[] | null;
context: string[] | null;
toolsCalled: ToolCall[] | null;
expectedTools: ToolCall[] | null;
metricsData: MetricData[];
annotations: AnnotationSummary[];
}uuidstringRequired
This is the unique identifier of the span.
Example: "<SPAN-UUID>"
traceUuidstringRequired
This is the uuid of the trace containing the span.
Example: "<TRACE-UUID>"
parentUuidstring | nullRequired
This is the uuid of the parent span, or null for a root span.
Example: "<PARENT-SPAN-UUID>"
namestring | nullRequired
This is the name of the span.
Example: "OpenAI Call"
typeSpanTypeRequired
See SpanType.
statusTraceSpanStatusRequired
See TraceSpanStatus.
startTimestringRequired
This is the time the span started.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the span ended.
Example: "2025-01-15T10:30:02+00:00"
errorstring | nullRequired
This is the error string that caused the span to fail, or null when no error occurred.
integrationstring | nullRequired
This is the integration associated with the span.
Example: "LangChain"
providerstring | nullRequired
This is the LLM provider used in an LLM span.
Example: "OpenAI"
modelstring | nullRequired
This is the LLM model used in an LLM span.
Example: "gpt-4o"
endpointstring | nullRequired
This is the API endpoint the model was called through in an LLM span.
costnumber | nullRequired
This is the total cost of the span in USD, or null when it is not known.
Example: 0.00018
inputTokenCostnumber | nullRequired
This is the total cost of the input tokens passed to the LLM model in an LLM span.
Example: 0.00006
outputTokenCostnumber | nullRequired
This is the total cost of the output tokens generated by the LLM model in an LLM span.
Example: 0.00012
costPerInputTokennumber | nullRequired
This is the cost per input token of the LLM model for an LLM span.
Example: 0.0000025
costPerOutputTokennumber | nullRequired
This is the cost per output token of the LLM model for an LLM span.
Example: 0.00001
inputTokenCountnumber | nullRequired
This is the total number of input tokens passed to the LLM model in an LLM span.
Example: 24
outputTokenCountnumber | nullRequired
This is the total number of output tokens generated by the LLM model in an LLM span.
Example: 12
promptAliasstring | nullRequired
This is the alias of your prompt which is stored on Confident AI.
Example: "geography-assistant"
promptVersionstring | nullRequired
This is the version assigned to your prompt on Confident AI.
Example: "00.00.01"
promptLabelstring | nullRequired
This is the label assigned to a specific version of prompt on the Confident AI platform.
Example: "production"
promptCommitHashstring | nullRequired
This is the hash of the current prompt being logged in the llm span.
Example: "bab04ce"
embedderstring | nullRequired
This is the embedder model used in a retriever span.
topKnumber | nullRequired
This is the top K chunks retrieved from your knowledge base in a retriever span.
chunkSizenumber | nullRequired
This is the chunk size of each retrieved context for a retriever span.
descriptionstring | nullRequired
This is a description if the span is a tool span.
agentHandoffsstring[] | nullRequired
This is the list of agent handoffs associated with an agent span.
availableToolsstring[] | nullRequired
This is the list of available tools associated with an agent span.
metadataRecord<string, unknown> | nullRequired
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metricCollectionNamestring | nullRequired
This is the name of the metric collection to evaluate the span.
Example: "LLM Collection Name"
inputstring | nullRequired
This is the input to the span. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputstring | nullRequired
This is the output of the span. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expectedOutputstring | nullRequired
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
retrievalContextstring[] | nullRequired
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextstring[] | nullRequired
This is the ideal retrieval context of your span, which is to be used for evaluation.
toolsCalledToolCall[] | nullRequired
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[] | nullRequired
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
metricsDataMetricData[]Required
This is the metrics data associated with the span.
See MetricData.
annotationsAnnotationSummary[]Required
This is the list of annotations associated with the span.
See AnnotationSummary.
SpanRequest
A span to record inside the trace. Its type selects the variant: omit it or send SPAN for a plain span, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, or AGENT for an agent step. Each variant accepts only its own fields.
SpanRequest = Union[
BaseSpanRequest,
LlmSpanRequest,
RetrieverSpanRequest,
ToolSpanRequest,
AgentSpanRequest,
]type SpanRequest =
| BaseSpanRequest
| LlmSpanRequest
| RetrieverSpanRequest
| ToolSpanRequest
| AgentSpanRequest;A SpanRequest is one of the shapes below. Send the fields of one of them, never a mix of both.
A plain span with no model, retriever, tool or agent detail.
class BaseSpanRequest:
type: Optional[Literal["SPAN"]] = None
uuid: str
name: str
input: Optional[Any] = None
output: Optional[Any] = None
error: Optional[str] = None
status: Optional[TraceSpanStatus] = None
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
parent_uuid: Optional[str] = Field(default=None, alias="parentUuid")
metadata: Optional[Dict[str, Any]] = None
metric_collection: Optional[str] = Field(default=None, alias="metricCollection")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
context: Optional[List[str]] = None
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
integration: Optional[str] = None
metrics_data: Optional[List[MetricDataConfig]] = Field(default=None, alias="metricsData")typeOptional[Literal["SPAN"]]
The type of the span. Omit it, or send SPAN, for a plain span.
Example: "SPAN"
uuidstrRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestrRequired
This is the name of the span.
Example: "OpenAI Call"
inputOptional[Any]
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputOptional[Any]
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorOptional[str]
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusOptional[TraceSpanStatus]
See TraceSpanStatus.
start_timestrRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parent_uuidOptional[str]
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataOptional[Dict[str, Any]]
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metric_collectionOptional[str]
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrieval_contextOptional[List[str]]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expected_outputOptional[str]
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
tools_calledOptional[List[ToolCall]]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationOptional[str]
This is the integration associated with the span.
Example: "LangChain"
metrics_dataOptional[List[MetricDataConfig]]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
interface BaseSpanRequest {
type?: "SPAN";
uuid: string;
name: string;
input?: unknown;
output?: unknown;
error?: string;
status?: TraceSpanStatus;
startTime: string;
endTime: string;
parentUuid?: string;
metadata?: Record<string, unknown>;
metricCollection?: string;
retrievalContext?: string[];
context?: string[];
expectedOutput?: string;
toolsCalled?: ToolCall[];
expectedTools?: ToolCall[];
integration?: string;
metricsData?: MetricDataConfig[];
}type"SPAN"
The type of the span. Omit it, or send SPAN, for a plain span.
Example: "SPAN"
uuidstringRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestringRequired
This is the name of the span.
Example: "OpenAI Call"
inputunknown
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputunknown
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorstring
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusTraceSpanStatus
See TraceSpanStatus.
startTimestringRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parentUuidstring
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataRecord<string, unknown>
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metricCollectionstring
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrievalContextstring[]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextstring[]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expectedOutputstring
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
toolsCalledToolCall[]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationstring
This is the integration associated with the span.
Example: "LangChain"
metricsDataMetricDataConfig[]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
A span recording a call to a language model, with its model, token counts, costs and the prompt it used.
class LlmSpanRequest:
type: Literal["LLM"]
uuid: str
name: str
input: Optional[Any] = None
output: Optional[Any] = None
error: Optional[str] = None
status: Optional[TraceSpanStatus] = None
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
parent_uuid: Optional[str] = Field(default=None, alias="parentUuid")
metadata: Optional[Dict[str, Any]] = None
metric_collection: Optional[str] = Field(default=None, alias="metricCollection")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
context: Optional[List[str]] = None
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
integration: Optional[str] = None
metrics_data: Optional[List[MetricDataConfig]] = Field(default=None, alias="metricsData")
model: Optional[str] = None
provider: Optional[str] = None
endpoint: Optional[str] = None
cost_per_input_token: Optional[float] = Field(default=None, alias="costPerInputToken")
cost_per_output_token: Optional[float] = Field(default=None, alias="costPerOutputToken")
input_token_count: Optional[int] = Field(default=None, alias="inputTokenCount")
output_token_count: Optional[int] = Field(default=None, alias="outputTokenCount")
prompt_alias: Optional[str] = Field(default=None, alias="promptAlias")
prompt_version: Optional[str] = Field(default=None, alias="promptVersion")
prompt_label: Optional[str] = Field(default=None, alias="promptLabel")
prompt_commit_hash: Optional[str] = Field(default=None, alias="promptCommitHash")typeLiteral["LLM"]Required
The type of the span, always LLM for an LLM span.
Example: "LLM"
uuidstrRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestrRequired
This is the name of the span.
Example: "OpenAI Call"
inputOptional[Any]
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputOptional[Any]
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorOptional[str]
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusOptional[TraceSpanStatus]
See TraceSpanStatus.
start_timestrRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parent_uuidOptional[str]
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataOptional[Dict[str, Any]]
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metric_collectionOptional[str]
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrieval_contextOptional[List[str]]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expected_outputOptional[str]
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
tools_calledOptional[List[ToolCall]]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationOptional[str]
This is the integration associated with the span.
Example: "LangChain"
metrics_dataOptional[List[MetricDataConfig]]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
modelOptional[str]
This is the LLM model used in the span.
Example: "gpt-4o"
providerOptional[str]
This is the provider of the generation model used in the span.
Example: "OpenAI"
endpointOptional[str]
This is the API endpoint the model was called through, for providers that expose more than one.
Example: "https://api.openai.com/v1/chat/completions"
cost_per_input_tokenOptional[float]
This is the cost per input token of the LLM model.
Example: 0.0000025
cost_per_output_tokenOptional[float]
This is the cost per output token of the LLM model.
Example: 0.00001
input_token_countOptional[int]
This is the number of input tokens passed to the LLM model.
Example: 24
output_token_countOptional[int]
This is the number of output tokens generated by the LLM model.
Example: 12
prompt_aliasOptional[str]
This is the alias of your prompt which is stored on Confident AI.
Example: "geography-assistant"
prompt_versionOptional[str]
This is the version assigned to your prompt on Confident AI.
Example: "00.00.01"
prompt_labelOptional[str]
This is the label assigned to a specific version of prompt on the Confident AI platform.
Example: "production"
prompt_commit_hashOptional[str]
This is the hash of the current prompt being logged in the llm span.
Example: "bab04ce"
interface LlmSpanRequest {
type: "LLM";
uuid: string;
name: string;
input?: unknown;
output?: unknown;
error?: string;
status?: TraceSpanStatus;
startTime: string;
endTime: string;
parentUuid?: string;
metadata?: Record<string, unknown>;
metricCollection?: string;
retrievalContext?: string[];
context?: string[];
expectedOutput?: string;
toolsCalled?: ToolCall[];
expectedTools?: ToolCall[];
integration?: string;
metricsData?: MetricDataConfig[];
model?: string;
provider?: string;
endpoint?: string;
costPerInputToken?: number;
costPerOutputToken?: number;
inputTokenCount?: number;
outputTokenCount?: number;
promptAlias?: string;
promptVersion?: string;
promptLabel?: string;
promptCommitHash?: string;
}type"LLM"Required
The type of the span, always LLM for an LLM span.
Example: "LLM"
uuidstringRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestringRequired
This is the name of the span.
Example: "OpenAI Call"
inputunknown
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputunknown
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorstring
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusTraceSpanStatus
See TraceSpanStatus.
startTimestringRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parentUuidstring
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataRecord<string, unknown>
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metricCollectionstring
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrievalContextstring[]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextstring[]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expectedOutputstring
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
toolsCalledToolCall[]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationstring
This is the integration associated with the span.
Example: "LangChain"
metricsDataMetricDataConfig[]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
modelstring
This is the LLM model used in the span.
Example: "gpt-4o"
providerstring
This is the provider of the generation model used in the span.
Example: "OpenAI"
endpointstring
This is the API endpoint the model was called through, for providers that expose more than one.
Example: "https://api.openai.com/v1/chat/completions"
costPerInputTokennumber
This is the cost per input token of the LLM model.
Example: 0.0000025
costPerOutputTokennumber
This is the cost per output token of the LLM model.
Example: 0.00001
inputTokenCountnumber
This is the number of input tokens passed to the LLM model.
Example: 24
outputTokenCountnumber
This is the number of output tokens generated by the LLM model.
Example: 12
promptAliasstring
This is the alias of your prompt which is stored on Confident AI.
Example: "geography-assistant"
promptVersionstring
This is the version assigned to your prompt on Confident AI.
Example: "00.00.01"
promptLabelstring
This is the label assigned to a specific version of prompt on the Confident AI platform.
Example: "production"
promptCommitHashstring
This is the hash of the current prompt being logged in the llm span.
Example: "bab04ce"
A span recording a knowledge-base lookup, with the embedder and retrieval settings it used.
class RetrieverSpanRequest:
type: Literal["RETRIEVER"]
uuid: str
name: str
input: Optional[Any] = None
output: Optional[Any] = None
error: Optional[str] = None
status: Optional[TraceSpanStatus] = None
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
parent_uuid: Optional[str] = Field(default=None, alias="parentUuid")
metadata: Optional[Dict[str, Any]] = None
metric_collection: Optional[str] = Field(default=None, alias="metricCollection")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
context: Optional[List[str]] = None
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
integration: Optional[str] = None
metrics_data: Optional[List[MetricDataConfig]] = Field(default=None, alias="metricsData")
embedder: str
top_k: Optional[int] = Field(default=None, alias="topK")
chunk_size: Optional[int] = Field(default=None, alias="chunkSize")typeLiteral["RETRIEVER"]Required
The type of the span, always RETRIEVER for a retriever span.
Example: "RETRIEVER"
uuidstrRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestrRequired
This is the name of the span.
Example: "OpenAI Call"
inputOptional[Any]
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputOptional[Any]
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorOptional[str]
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusOptional[TraceSpanStatus]
See TraceSpanStatus.
start_timestrRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parent_uuidOptional[str]
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataOptional[Dict[str, Any]]
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metric_collectionOptional[str]
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrieval_contextOptional[List[str]]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expected_outputOptional[str]
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
tools_calledOptional[List[ToolCall]]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationOptional[str]
This is the integration associated with the span.
Example: "LangChain"
metrics_dataOptional[List[MetricDataConfig]]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
embedderstrRequired
This is the embedder model used in the span.
Example: "text-embedding-3-small"
top_kOptional[int]
This is the top K chunks retrieved from your knowledge base.
Example: 3
chunk_sizeOptional[int]
This is the chunk size of each retrieved context.
Example: 512
interface RetrieverSpanRequest {
type: "RETRIEVER";
uuid: string;
name: string;
input?: unknown;
output?: unknown;
error?: string;
status?: TraceSpanStatus;
startTime: string;
endTime: string;
parentUuid?: string;
metadata?: Record<string, unknown>;
metricCollection?: string;
retrievalContext?: string[];
context?: string[];
expectedOutput?: string;
toolsCalled?: ToolCall[];
expectedTools?: ToolCall[];
integration?: string;
metricsData?: MetricDataConfig[];
embedder: string;
topK?: number;
chunkSize?: number;
}type"RETRIEVER"Required
The type of the span, always RETRIEVER for a retriever span.
Example: "RETRIEVER"
uuidstringRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestringRequired
This is the name of the span.
Example: "OpenAI Call"
inputunknown
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputunknown
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorstring
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusTraceSpanStatus
See TraceSpanStatus.
startTimestringRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parentUuidstring
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataRecord<string, unknown>
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metricCollectionstring
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrievalContextstring[]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextstring[]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expectedOutputstring
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
toolsCalledToolCall[]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationstring
This is the integration associated with the span.
Example: "LangChain"
metricsDataMetricDataConfig[]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
embedderstringRequired
This is the embedder model used in the span.
Example: "text-embedding-3-small"
topKnumber
This is the top K chunks retrieved from your knowledge base.
Example: 3
chunkSizenumber
This is the chunk size of each retrieved context.
Example: 512
A span recording a tool call.
class ToolSpanRequest:
type: Literal["TOOL"]
uuid: str
name: str
input: Optional[Any] = None
output: Optional[Any] = None
error: Optional[str] = None
status: Optional[TraceSpanStatus] = None
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
parent_uuid: Optional[str] = Field(default=None, alias="parentUuid")
metadata: Optional[Dict[str, Any]] = None
metric_collection: Optional[str] = Field(default=None, alias="metricCollection")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
context: Optional[List[str]] = None
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
integration: Optional[str] = None
metrics_data: Optional[List[MetricDataConfig]] = Field(default=None, alias="metricsData")
description: Optional[str] = NonetypeLiteral["TOOL"]Required
The type of the span, always TOOL for a tool span.
Example: "TOOL"
uuidstrRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestrRequired
This is the name of the span.
Example: "OpenAI Call"
inputOptional[Any]
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputOptional[Any]
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorOptional[str]
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusOptional[TraceSpanStatus]
See TraceSpanStatus.
start_timestrRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parent_uuidOptional[str]
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataOptional[Dict[str, Any]]
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metric_collectionOptional[str]
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrieval_contextOptional[List[str]]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expected_outputOptional[str]
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
tools_calledOptional[List[ToolCall]]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationOptional[str]
This is the integration associated with the span.
Example: "LangChain"
metrics_dataOptional[List[MetricDataConfig]]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
descriptionOptional[str]
This is the description of the tool used in the span.
Example: "Looks up the capital city of a country."
interface ToolSpanRequest {
type: "TOOL";
uuid: string;
name: string;
input?: unknown;
output?: unknown;
error?: string;
status?: TraceSpanStatus;
startTime: string;
endTime: string;
parentUuid?: string;
metadata?: Record<string, unknown>;
metricCollection?: string;
retrievalContext?: string[];
context?: string[];
expectedOutput?: string;
toolsCalled?: ToolCall[];
expectedTools?: ToolCall[];
integration?: string;
metricsData?: MetricDataConfig[];
description?: string;
}type"TOOL"Required
The type of the span, always TOOL for a tool span.
Example: "TOOL"
uuidstringRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestringRequired
This is the name of the span.
Example: "OpenAI Call"
inputunknown
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputunknown
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorstring
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusTraceSpanStatus
See TraceSpanStatus.
startTimestringRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parentUuidstring
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataRecord<string, unknown>
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metricCollectionstring
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrievalContextstring[]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextstring[]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expectedOutputstring
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
toolsCalledToolCall[]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationstring
This is the integration associated with the span.
Example: "LangChain"
metricsDataMetricDataConfig[]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
descriptionstring
This is the description of the tool used in the span.
Example: "Looks up the capital city of a country."
A span recording an agent step, with the tools and handoffs available to it.
class AgentSpanRequest:
type: Literal["AGENT"]
uuid: str
name: str
input: Optional[Any] = None
output: Optional[Any] = None
error: Optional[str] = None
status: Optional[TraceSpanStatus] = None
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
parent_uuid: Optional[str] = Field(default=None, alias="parentUuid")
metadata: Optional[Dict[str, Any]] = None
metric_collection: Optional[str] = Field(default=None, alias="metricCollection")
retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
context: Optional[List[str]] = None
expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
integration: Optional[str] = None
metrics_data: Optional[List[MetricDataConfig]] = Field(default=None, alias="metricsData")
available_tools: List[str] = Field(alias="availableTools")
agent_handoffs: List[str] = Field(alias="agentHandoffs")typeLiteral["AGENT"]Required
The type of the span, always AGENT for an agent span.
Example: "AGENT"
uuidstrRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestrRequired
This is the name of the span.
Example: "OpenAI Call"
inputOptional[Any]
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputOptional[Any]
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorOptional[str]
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusOptional[TraceSpanStatus]
See TraceSpanStatus.
start_timestrRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parent_uuidOptional[str]
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataOptional[Dict[str, Any]]
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metric_collectionOptional[str]
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrieval_contextOptional[List[str]]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expected_outputOptional[str]
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
tools_calledOptional[List[ToolCall]]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationOptional[str]
This is the integration associated with the span.
Example: "LangChain"
metrics_dataOptional[List[MetricDataConfig]]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
available_toolsList[str]Required
This is the list of names of available tools to be used in the span.
Example: ["get_capital"]
agent_handoffsList[str]Required
This is the list of potential agent handoffs in the span.
Example: ["geography_agent"]
interface AgentSpanRequest {
type: "AGENT";
uuid: string;
name: string;
input?: unknown;
output?: unknown;
error?: string;
status?: TraceSpanStatus;
startTime: string;
endTime: string;
parentUuid?: string;
metadata?: Record<string, unknown>;
metricCollection?: string;
retrievalContext?: string[];
context?: string[];
expectedOutput?: string;
toolsCalled?: ToolCall[];
expectedTools?: ToolCall[];
integration?: string;
metricsData?: MetricDataConfig[];
availableTools: string[];
agentHandoffs: string[];
}type"AGENT"Required
The type of the span, always AGENT for an agent span.
Example: "AGENT"
uuidstringRequired
The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.
Example: "<SPAN-UUID>"
namestringRequired
This is the name of the span.
Example: "OpenAI Call"
inputunknown
This is the input to the span, as a string or any JSON value.
Example: "What is the capital of France?"
outputunknown
This is the output of the span, as a string or any JSON value.
Example: "The capital of France is Paris."
errorstring
This is the error message, if an error occurred inside the span.
Example: "The model timed out after 30 seconds."
statusTraceSpanStatus
See TraceSpanStatus.
startTimestringRequired
This is the time the span started, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the span ended, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:02+00:00"
parentUuidstring
This is the unique identifier of the span's parent span. Omit it for a root span.
Example: "<PARENT-SPAN-UUID>"
metadataRecord<string, unknown>
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metricCollectionstring
This is the metric collection to be used for evaluating the span.
Example: "LLM Collection Name"
retrievalContextstring[]
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextstring[]
This is the ideal retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital of France."]
expectedOutputstring
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
toolsCalledToolCall[]
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[]
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
integrationstring
This is the integration associated with the span.
Example: "LangChain"
metricsDataMetricDataConfig[]
Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.
See MetricDataConfig.
availableToolsstring[]Required
This is the list of names of available tools to be used in the span.
Example: ["get_capital"]
agentHandoffsstring[]Required
This is the list of potential agent handoffs in the span.
Example: ["geography_agent"]
SpanType
The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.
class SpanType(Enum):
SPAN = "SPAN"
AGENT = "AGENT"
TOOL = "TOOL"
RETRIEVER = "RETRIEVER"
LLM = "LLM"enum SpanType {
SPAN = "SPAN",
AGENT = "AGENT",
TOOL = "TOOL",
RETRIEVER = "RETRIEVER",
LLM = "LLM",
}SPAN · AGENT · TOOL · RETRIEVER · LLM
ThreadRequest
Thread-level fields applied to the thread record. id is an alternate way to specify the thread and must match top-level threadId if both are provided. metadata and tags only take effect when a thread id is resolvable; successive ingestions merge metadata keys, while tags replace any prior value.
class ThreadRequest:
id: Optional[str] = None
metadata: Optional[Dict[str, Any]] = None
tags: Optional[List[str]] = NoneidOptional[str]
The thread id. Equivalent to top-level threadId; if both are set they must match.
Example: "thread-42"
metadataOptional[Dict[str, Any]]
Custom key/value metadata to attach to the thread. Values can be any JSON-serializable type and are stringified server-side. Successive ingestions for the same thread merge metadata keys.
Example: {"client":"acme-corp","agentId":"geography-agent"}
tagsOptional[List[str]]
Tags to set on the thread. Replaces any previously stored tags.
Example: ["vip"]
interface ThreadRequest {
id?: string;
metadata?: Record<string, unknown> | null;
tags?: string[] | null;
}idstring
The thread id. Equivalent to top-level threadId; if both are set they must match.
Example: "thread-42"
metadataRecord<string, unknown> | null
Custom key/value metadata to attach to the thread. Values can be any JSON-serializable type and are stringified server-side. Successive ingestions for the same thread merge metadata keys.
Example: {"client":"acme-corp","agentId":"geography-agent"}
tagsstring[] | null
Tags to set on the thread. Replaces any previously stored tags.
Example: ["vip"]
ToolCall
A tool your LLM application invoked, with what it passed in and what came back.
class ToolCall:
name: str
type: Optional[ToolCallType] = None
description: Optional[str] = None
input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
output: Optional[Any] = None
reasoning: Optional[str] = NonenamestrRequired
This is the name of the tool.
Example: "get_landmark_info"
typeOptional[ToolCallType]
See ToolCallType.
descriptionOptional[str]
This is the description of the tool.
Example: "This tool gives information about a mountain."
input_parametersOptional[Dict[str, Any]]
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputOptional[Any]
This is the output of the tool.
Example: "8,848 metres"
reasoningOptional[str]
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
interface ToolCall {
name: string;
type?: ToolCallType;
description?: string;
inputParameters?: Record<string, unknown> | null;
output?: unknown;
reasoning?: string;
}namestringRequired
This is the name of the tool.
Example: "get_landmark_info"
typeToolCallType
See ToolCallType.
descriptionstring
This is the description of the tool.
Example: "This tool gives information about a mountain."
inputParametersRecord<string, unknown> | null
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputunknown
This is the output of the tool.
Example: "8,848 metres"
reasoningstring
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
ToolCallType
The type of the tool call, either a function or an MCP tool.
class ToolCallType(Enum):
FUNCTION = "FUNCTION"
MCP = "MCP"enum ToolCallType {
FUNCTION = "FUNCTION",
MCP = "MCP",
}FUNCTION · MCP
Trace
A trace with its full input and output, evaluation fields, classifier labels, results and annotations, and its spans when retrieved by id.
class Trace:
uuid: str
name: Optional[str]
status: TraceSpanStatus
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
latency: int
cost: Optional[float]
thread_id: Optional[str] = Field(alias="threadId")
user_id: Optional[str] = Field(alias="userId")
customer_id: Optional[str] = Field(alias="customerId")
environment: Environment
tags: Optional[List[str]]
metadata: Optional[Dict[str, Any]]
input: Optional[str]
output: Optional[str]
expected_output: Optional[str] = Field(alias="expectedOutput")
retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
context: Optional[List[str]]
tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
test_case_id: Optional[str] = Field(alias="testCaseId")
metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
labels: Dict[str, Classification]
spans: Optional[List[Span]] = None
metrics_data: List[MetricData] = Field(alias="metricsData")
annotations: List[AnnotationSummary]uuidstrRequired
This is the unique identifier of the trace.
Example: "<TRACE-UUID>"
nameOptional[str]Required
This is the name of the trace.
Example: "Geography QA"
statusTraceSpanStatusRequired
See TraceSpanStatus.
start_timestrRequired
This is the time the trace started.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the trace ended.
Example: "2025-01-15T10:30:05+00:00"
latencyintRequired
This is how long the trace took, in milliseconds.
Example: 5000
costOptional[float]Required
This is the total cost of the trace in USD, summed from its spans, or null when it is not known.
Example: 0.00018
thread_idOptional[str]Required
This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.
Example: "thread-42"
user_idOptional[str]Required
This is the user id you provided for this trace, or null when you did not.
Example: "end-user-42"
customer_idOptional[str]Required
This is the customer id you provided for this trace, or null when you did not.
Example: "acme-hotels"
environmentEnvironmentRequired
See Environment.
tagsOptional[List[str]]Required
This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.
Example: ["geography"]
metadataOptional[Dict[str, Any]]Required
This is any additional metadata associated with the trace.
Example: {"client":"acme-corp"}
inputOptional[str]Required
This is the input to the trace. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputOptional[str]Required
This is the output of the trace. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expected_outputOptional[str]Required
This is the expected output associated with the trace, to be used for evaluations.
Example: "Paris"
retrieval_contextOptional[List[str]]Required
This is the retrieval context associated with the trace, to be used for evaluations.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]Required
This is the ideal retrieval context associated with the trace, to be used for evaluations.
tools_calledOptional[List[ToolCall]]Required
This is the list of tools called by the trace, to be used for evaluations.
See ToolCall.
expected_toolsOptional[List[ToolCall]]Required
This is the list of expected tools associated with the trace, to be used for evaluations.
See ToolCall.
test_case_idOptional[str]Required
This is the test case id of the trace, which is only set if the trace was created while evaluating a test case.
metric_collection_nameOptional[str]Required
This is the name of the metric collection assigned to evaluate the trace.
Example: "Collection Name"
labelsDict[str, Classification]Required
The labels your project's classifiers assigned to the trace, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital city of a country."}}
spansOptional[List[Span]]
This is the list of spans in the trace, present when the trace is retrieved by id. A thread's traces omit their spans.
See Span.
metrics_dataList[MetricData]Required
This is the list of metrics data associated with the trace after running evaluations.
See MetricData.
annotationsList[AnnotationSummary]Required
This is the list of annotations associated with the trace.
See AnnotationSummary.
interface Trace {
uuid: string;
name: string | null;
status: TraceSpanStatus;
startTime: string;
endTime: string;
latency: number;
cost: number | null;
threadId: string | null;
userId: string | null;
customerId: string | null;
environment: Environment;
tags: string[] | null;
metadata: Record<string, unknown> | null;
input: string | null;
output: string | null;
expectedOutput: string | null;
retrievalContext: string[] | null;
context: string[] | null;
toolsCalled: ToolCall[] | null;
expectedTools: ToolCall[] | null;
testCaseId: string | null;
metricCollectionName: string | null;
labels: Record<string, Classification>;
spans?: Span[];
metricsData: MetricData[];
annotations: AnnotationSummary[];
}uuidstringRequired
This is the unique identifier of the trace.
Example: "<TRACE-UUID>"
namestring | nullRequired
This is the name of the trace.
Example: "Geography QA"
statusTraceSpanStatusRequired
See TraceSpanStatus.
startTimestringRequired
This is the time the trace started.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the trace ended.
Example: "2025-01-15T10:30:05+00:00"
latencynumberRequired
This is how long the trace took, in milliseconds.
Example: 5000
costnumber | nullRequired
This is the total cost of the trace in USD, summed from its spans, or null when it is not known.
Example: 0.00018
threadIdstring | nullRequired
This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.
Example: "thread-42"
userIdstring | nullRequired
This is the user id you provided for this trace, or null when you did not.
Example: "end-user-42"
customerIdstring | nullRequired
This is the customer id you provided for this trace, or null when you did not.
Example: "acme-hotels"
environmentEnvironmentRequired
See Environment.
tagsstring[] | nullRequired
This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.
Example: ["geography"]
metadataRecord<string, unknown> | nullRequired
This is any additional metadata associated with the trace.
Example: {"client":"acme-corp"}
inputstring | nullRequired
This is the input to the trace. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputstring | nullRequired
This is the output of the trace. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expectedOutputstring | nullRequired
This is the expected output associated with the trace, to be used for evaluations.
Example: "Paris"
retrievalContextstring[] | nullRequired
This is the retrieval context associated with the trace, to be used for evaluations.
Example: ["Paris is the capital and most populous city of France."]
contextstring[] | nullRequired
This is the ideal retrieval context associated with the trace, to be used for evaluations.
toolsCalledToolCall[] | nullRequired
This is the list of tools called by the trace, to be used for evaluations.
See ToolCall.
expectedToolsToolCall[] | nullRequired
This is the list of expected tools associated with the trace, to be used for evaluations.
See ToolCall.
testCaseIdstring | nullRequired
This is the test case id of the trace, which is only set if the trace was created while evaluating a test case.
metricCollectionNamestring | nullRequired
This is the name of the metric collection assigned to evaluate the trace.
Example: "Collection Name"
labelsRecord<string, Classification>Required
The labels your project's classifiers assigned to the trace, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital city of a country."}}
spansSpan[]
This is the list of spans in the trace, present when the trace is retrieved by id. A thread's traces omit their spans.
See Span.
metricsDataMetricData[]Required
This is the list of metrics data associated with the trace after running evaluations.
See MetricData.
annotationsAnnotationSummary[]Required
This is the list of annotations associated with the trace.
See AnnotationSummary.
TraceAttachment
Payload for a multimodal attachment referenced by id in [DEEPEVAL:IMAGE:…] or [DEEPEVAL:PDF:…] markers. Provide either a url, or dataBase64 with mimeType.
class TraceAttachment:
url: Optional[str] = None
data_base64: Optional[str] = Field(default=None, alias="dataBase64")
mime_type: Optional[str] = Field(default=None, alias="mimeType")urlOptional[str]
Public URL of the attachment. Send either url, or dataBase64 with mimeType, not both.
Example: "https://example.com/paris.pdf"
data_base64Optional[str]
Base64-encoded file bytes, as an alternative to url.
Example: "JVBERi0xLjQK"
mime_typeOptional[str]
MIME type of the attachment, required when using dataBase64.
Example: "application/pdf"
interface TraceAttachment {
url?: string;
dataBase64?: string;
mimeType?: string;
}urlstring
Public URL of the attachment. Send either url, or dataBase64 with mimeType, not both.
Example: "https://example.com/paris.pdf"
dataBase64string
Base64-encoded file bytes, as an alternative to url.
Example: "JVBERi0xLjQK"
mimeTypestring
MIME type of the attachment, required when using dataBase64.
Example: "application/pdf"
TraceList
class TraceList:
traces: List[TraceSummary]
total_traces: Optional[int] = Field(default=None, alias="totalTraces")
next_cursor: Optional[str] = Field(alias="nextCursor")tracesList[TraceSummary]Required
This is the list of traces for the current page.
See TraceSummary.
total_tracesOptional[int]
This is the total number of traces matching the query across all pages. Present on the first page only; omitted when a cursor is given.
Example: 1
next_cursorOptional[str]Required
The value to pass as cursor to get the next page, or null when this is the last page.
interface TraceList {
traces: TraceSummary[];
totalTraces?: number;
nextCursor: string | null;
}tracesTraceSummary[]Required
This is the list of traces for the current page.
See TraceSummary.
totalTracesnumber
This is the total number of traces matching the query across all pages. Present on the first page only; omitted when a cursor is given.
Example: 1
nextCursorstring | nullRequired
The value to pass as cursor to get the next page, or null when this is the last page.
TraceRef
class TraceRef:
uuid: struuidstrRequired
This is the uuid of the trace. It is the uuid you sent, or its UUID hash when the value you sent was not a UUID.
Example: "<TRACE-UUID>"
interface TraceRef {
uuid: string;
}uuidstringRequired
This is the uuid of the trace. It is the uuid you sent, or its UUID hash when the value you sent was not a UUID.
Example: "<TRACE-UUID>"
TraceSortBy
The trace field to sort by: createdAt orders by the trace's start time and endedAt by its end time.
class TraceSortBy(Enum):
CREATEDAT = "createdAt"
ENDEDAT = "endedAt"enum TraceSortBy {
CREATEDAT = "createdAt",
ENDEDAT = "endedAt",
}CREATEDAT · ENDEDAT
TraceSpanStatus
This represents the error status of a trace or span: SUCCESS when it completed, ERRORED when it failed.
class TraceSpanStatus(Enum):
SUCCESS = "SUCCESS"
ERRORED = "ERRORED"enum TraceSpanStatus {
SUCCESS = "SUCCESS",
ERRORED = "ERRORED",
}SUCCESS · ERRORED
TraceSummary
A trace as it appears in a list: its timing, cost, thread and user with a preview of its input and output, but without its spans, evaluation fields, results or annotations.
class TraceSummary:
uuid: str
name: Optional[str]
status: TraceSpanStatus
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
latency: int
cost: Optional[float]
thread_id: Optional[str] = Field(alias="threadId")
user_id: Optional[str] = Field(alias="userId")
customer_id: Optional[str] = Field(alias="customerId")
environment: Environment
tags: Optional[List[str]]
metadata: Optional[Dict[str, Any]]
input_preview: Optional[str] = Field(alias="inputPreview")
output_preview: Optional[str] = Field(alias="outputPreview")uuidstrRequired
This is the unique identifier of the trace.
Example: "<TRACE-UUID>"
nameOptional[str]Required
This is the name of the trace.
Example: "Geography QA"
statusTraceSpanStatusRequired
See TraceSpanStatus.
start_timestrRequired
This is the time the trace started.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the trace ended.
Example: "2025-01-15T10:30:05+00:00"
latencyintRequired
This is how long the trace took, in milliseconds.
Example: 5000
costOptional[float]Required
This is the total cost of the trace in USD, summed from its spans, or null when it is not known.
Example: 0.00018
thread_idOptional[str]Required
This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.
Example: "thread-42"
user_idOptional[str]Required
This is the user id you provided for this trace, or null when you did not.
Example: "end-user-42"
customer_idOptional[str]Required
This is the customer id you provided for this trace, or null when you did not.
Example: "acme-hotels"
environmentEnvironmentRequired
See Environment.
tagsOptional[List[str]]Required
This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.
Example: ["geography"]
metadataOptional[Dict[str, Any]]Required
This is any additional metadata associated with the trace.
Example: {"client":"acme-corp"}
input_previewOptional[str]Required
The first characters of the trace's input, or null when it has none. Retrieve the trace by id for the full value.
Example: "What is the capital of France?"
output_previewOptional[str]Required
The first characters of the trace's output, or null when it has none. Retrieve the trace by id for the full value.
Example: "The capital of France is Paris."
interface TraceSummary {
uuid: string;
name: string | null;
status: TraceSpanStatus;
startTime: string;
endTime: string;
latency: number;
cost: number | null;
threadId: string | null;
userId: string | null;
customerId: string | null;
environment: Environment;
tags: string[] | null;
metadata: Record<string, unknown> | null;
inputPreview: string | null;
outputPreview: string | null;
}uuidstringRequired
This is the unique identifier of the trace.
Example: "<TRACE-UUID>"
namestring | nullRequired
This is the name of the trace.
Example: "Geography QA"
statusTraceSpanStatusRequired
See TraceSpanStatus.
startTimestringRequired
This is the time the trace started.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the trace ended.
Example: "2025-01-15T10:30:05+00:00"
latencynumberRequired
This is how long the trace took, in milliseconds.
Example: 5000
costnumber | nullRequired
This is the total cost of the trace in USD, summed from its spans, or null when it is not known.
Example: 0.00018
threadIdstring | nullRequired
This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.
Example: "thread-42"
userIdstring | nullRequired
This is the user id you provided for this trace, or null when you did not.
Example: "end-user-42"
customerIdstring | nullRequired
This is the customer id you provided for this trace, or null when you did not.
Example: "acme-hotels"
environmentEnvironmentRequired
See Environment.
tagsstring[] | nullRequired
This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.
Example: ["geography"]
metadataRecord<string, unknown> | nullRequired
This is any additional metadata associated with the trace.
Example: {"client":"acme-corp"}
inputPreviewstring | nullRequired
The first characters of the trace's input, or null when it has none. Retrieve the trace by id for the full value.
Example: "What is the capital of France?"
outputPreviewstring | nullRequired
The first characters of the trace's output, or null when it has none. Retrieve the trace by id for the full value.
Example: "The capital of France is Paris."
UserReference
A Confident AI user, as referenced by the records they created.
class UserReference:
id: str
email: str
name: Optional[str]
image: Optional[str]idstrRequired
This is the id of the user.
Example: "<USER-ID>"
emailstrRequired
This is the email address of the user.
Example: "jane@acme.com"
nameOptional[str]Required
This is the display name of the user, or null when they have not set one.
Example: "Jane Doe"
imageOptional[str]Required
This is the URL of the user's avatar, or null when they have none.
interface UserReference {
id: string;
email: string;
name: string | null;
image: string | null;
}idstringRequired
This is the id of the user.
Example: "<USER-ID>"
emailstringRequired
This is the email address of the user.
Example: "jane@acme.com"
namestring | nullRequired
This is the display name of the user, or null when they have not set one.
Example: "Jane Doe"
imagestring | nullRequired
This is the URL of the user's avatar, or null when they have none.
UserRequest
End-user-level fields applied to the end-user record. id is an alternate way to specify the end user and must match top-level userId if both are provided. name only takes effect when an end-user id is resolvable; successive ingestions take the latest name.
class UserRequest:
id: Optional[str] = None
name: Optional[str] = NoneidOptional[str]
The end-user id. Equivalent to top-level userId; if both are set they must match.
Example: "end-user-42"
nameOptional[str]
A human-readable display name for the end user, shown instead of the id.
Example: "Marta Ruiz"
interface UserRequest {
id?: string;
name?: string | null;
}idstring
The end-user id. Equivalent to top-level userId; if both are set they must match.
Example: "end-user-42"
namestring | null
A human-readable display name for the end user, shown instead of the id.
Example: "Marta Ruiz"
Last updated on