Threads
Overview
The Confident AI SDK exposes every Thread method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
List Threads
Lists the threads in your Confident AI project one page at a time, most recently active first by default. Each thread is returned as a summary; retrieve a thread by id for its traces, evaluation results and annotations.
from confident_ai import ConfidentAI
from confident_ai.common import Environment
from confident_ai.threads import ThreadSortBy
client = ConfidentAI()
result = client.threads.list(
page_size=25,
cursor="<NEXT-CURSOR>",
start="2025-01-01T00:00:00+00:00",
end="2025-01-31T23:59:59+00:00",
ascending="false",
sort_by=ThreadSortBy.LASTACTIVITY,
environment=Environment.PRODUCTION,
)For async mode, call a_list and await it as shown below:
result = await client.threads.a_list(...)Parameters
| Parameter | Type | Description |
|---|---|---|
page_size | Optional[int] | The number of results per page, at most 100. Defaults to 25. |
cursor | Optional[str] | This is used for pagination, and should be set to the nextCursor value returned in the previous response to get the next page of results. |
start | Optional[str] | This filters for results created at or after the specified start datetime, in ISO 8601 format. Defaults to 60 days ago. |
end | Optional[str] | This filters for results created before the specified end datetime, in ISO 8601 format. Defaults to the current time. |
ascending | Optional[Literal['true', 'false']] | This determines if the field specified in sortBy should be in ascending order. Defaults to false, which returns the newest results first. |
sort_by | Optional[ThreadSortBy] | This determines the field to sort by. Defaults to lastActivity. See ThreadSortBy. |
environment | Optional[Environment] | This filters the threads by the environment where their traces were created, and returns threads from all environments if not specified. See Environment. |
import { ConfidentAI } from "confident-ai";
import { Environment } from "confident-ai/common";
import { ThreadSortBy } from "confident-ai/threads";
const client = new ConfidentAI();
const result = await client.threads.list(
{
pageSize: 25,
cursor: "<NEXT-CURSOR>",
start: "2025-01-01T00:00:00+00:00",
end: "2025-01-31T23:59:59+00:00",
ascending: "false",
sortBy: ThreadSortBy.LASTACTIVITY,
environment: Environment.PRODUCTION
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
pageSize | number | The number of results per page, at most 100. Defaults to 25. |
cursor | string | This is used for pagination, and should be set to the nextCursor value returned in the previous response to get the next page of results. |
start | string | This filters for results created at or after the specified start datetime, in ISO 8601 format. Defaults to 60 days ago. |
end | string | This filters for results created before the specified end datetime, in ISO 8601 format. Defaults to the current time. |
ascending | "true" | "false" | This determines if the field specified in sortBy should be in ascending order. Defaults to false, which returns the newest results first. |
sortBy | ThreadSortBy | This determines the field to sort by. Defaults to lastActivity. See ThreadSortBy. |
environment | Environment | This filters the threads by the environment where their traces were created, and returns threads from all environments if not specified. See Environment. |
Returns
This method returns an object of type ThreadList.
Get Thread
Retrieves a thread by id from your Confident AI project, with its evaluation results, annotations and the first 100 traces of the conversation, oldest first.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.threads.get(thread_id="thread-42")For async mode, call a_get and await it as shown below:
result = await client.threads.a_get(...)Parameters
| Parameter | Type | Description |
|---|---|---|
thread_id | str | Required. The id of the thread, as you supplied it when creating its traces. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.threads.get("thread-42");Parameters
| Parameter | Type | Description |
|---|---|---|
threadId | string | Required. The id of the thread, as you supplied it when creating its traces. |
Returns
This method returns an object of type Thread.
Types
AnnotationFieldType
The kind of value an annotation holds: TEXT, NUMBER, FLOAT or BOOLEAN for a free value, CHOICE or MULTIPLE_CHOICE for a choice from the options in the field's config, and THUMBS_RATING or FIVE_STAR_RATING for a rating.
class AnnotationFieldType(Enum):
TEXT = "TEXT"
NUMBER = "NUMBER"
FLOAT = "FLOAT"
BOOLEAN = "BOOLEAN"
CHOICE = "CHOICE"
MULTIPLE_CHOICE = "MULTIPLE_CHOICE"
FIVE_STAR_RATING = "FIVE_STAR_RATING"
THUMBS_RATING = "THUMBS_RATING"enum AnnotationFieldType {
TEXT = "TEXT",
NUMBER = "NUMBER",
FLOAT = "FLOAT",
BOOLEAN = "BOOLEAN",
CHOICE = "CHOICE",
MULTIPLE_CHOICE = "MULTIPLE_CHOICE",
FIVE_STAR_RATING = "FIVE_STAR_RATING",
THUMBS_RATING = "THUMBS_RATING",
}TEXT · NUMBER · FLOAT · BOOLEAN · CHOICE · MULTIPLE_CHOICE · FIVE_STAR_RATING · THUMBS_RATING
AnnotationSummary
A human rating as it appears under the trace, span or thread it was left on, without repeating the ids of that target.
class AnnotationSummary:
id: str
field_type: AnnotationFieldType = Field(alias="fieldType")
value: Optional[Union[str, float, bool, List[str]]]
name: Optional[str]
explanation: Optional[str]
expected_outcome: Optional[str] = Field(alias="expectedOutcome")
expected_output: Optional[str] = Field(alias="expectedOutput")
created_at: str = Field(alias="createdAt")
user: Optional[UserReference]idstrRequired
This is the id of the annotation generated by Confident AI.
Example: "<ANNOTATION-ID>"
field_typeAnnotationFieldTypeRequired
See AnnotationFieldType.
valueOptional[Union[str, float, bool, List[str]]]Required
The annotated value, in the shape its fieldType expects: a boolean for THUMBS_RATING (true is thumbs up) and BOOLEAN, a whole number from 1 to 5 for FIVE_STAR_RATING, a number for NUMBER or FLOAT, a string for TEXT and CHOICE, and a list of strings for MULTIPLE_CHOICE. Null when the annotation was left without a value.
Example: true
nameOptional[str]Required
The name of the annotation.
explanationOptional[str]Required
This is the explanation for the annotation.
Example: "Correct and concise."
expected_outcomeOptional[str]Required
This is the annotated expected outcome, for conversation annotations.
expected_outputOptional[str]Required
This is the annotated expected output, for span and trace annotations.
Example: "The capital of France is Paris."
created_atstrRequired
The timestamp when the annotation was created.
Example: "2025-01-15T11:00:00+00:00"
userOptional[UserReference]Required
See UserReference.
interface AnnotationSummary {
id: string;
fieldType: AnnotationFieldType;
value: string | number | boolean | string[] | null;
name: string | null;
explanation: string | null;
expectedOutcome: string | null;
expectedOutput: string | null;
createdAt: string;
user: UserReference | null;
}idstringRequired
This is the id of the annotation generated by Confident AI.
Example: "<ANNOTATION-ID>"
fieldTypeAnnotationFieldTypeRequired
See AnnotationFieldType.
valuestring | number | boolean | string[] | nullRequired
The annotated value, in the shape its fieldType expects: a boolean for THUMBS_RATING (true is thumbs up) and BOOLEAN, a whole number from 1 to 5 for FIVE_STAR_RATING, a number for NUMBER or FLOAT, a string for TEXT and CHOICE, and a list of strings for MULTIPLE_CHOICE. Null when the annotation was left without a value.
Example: true
namestring | nullRequired
The name of the annotation.
explanationstring | nullRequired
This is the explanation for the annotation.
Example: "Correct and concise."
expectedOutcomestring | nullRequired
This is the annotated expected outcome, for conversation annotations.
expectedOutputstring | nullRequired
This is the annotated expected output, for span and trace annotations.
Example: "The capital of France is Paris."
createdAtstringRequired
The timestamp when the annotation was created.
Example: "2025-01-15T11:00:00+00:00"
userUserReference | nullRequired
See UserReference.
Classification
A label assigned by one of your project's classifiers, with the reason it was chosen.
class Classification:
label: str
reason: strlabelstrRequired
The label the classifier assigned.
Example: "geography"
reasonstrRequired
The classifier's reason for choosing the label.
Example: "The user asks for the capital city of a country."
interface Classification {
label: string;
reason: string;
}labelstringRequired
The label the classifier assigned.
Example: "geography"
reasonstringRequired
The classifier's reason for choosing the label.
Example: "The user asks for the capital city of a country."
Environment
This is the environment where your trace was posted, which helps with separating and debugging traces from different environments on the Confident AI platform.
class Environment(Enum):
PRODUCTION = "production"
DEVELOPMENT = "development"
STAGING = "staging"
TESTING = "testing"enum Environment {
PRODUCTION = "production",
DEVELOPMENT = "development",
STAGING = "staging",
TESTING = "testing",
}PRODUCTION · DEVELOPMENT · STAGING · TESTING
EvaluationErrorType
Why an evaluation errored: the AI connection or a transformer failed, the evaluation model failed, the test case lacked the parameters the metric needs, or an internal error occurred.
class EvaluationErrorType(Enum):
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR"
TRANSFORMER_ERROR = "TRANSFORMER_ERROR"
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR"
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS"
INTERNAL_ERROR = "INTERNAL_ERROR"enum EvaluationErrorType {
AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR",
TRANSFORMER_ERROR = "TRANSFORMER_ERROR",
EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR",
INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS",
INTERNAL_ERROR = "INTERNAL_ERROR",
}AI_CONNECTION_ERROR · TRANSFORMER_ERROR · EVALUATION_MODEL_ERROR · INVALID_TEST_CASE_PARAMETERS · INTERNAL_ERROR
MetricData
The result of an evaluated metric, with the ids of whatever it was recorded against.
class MetricData:
id: str
name: str
score: Optional[float]
reason: Optional[str]
success: Optional[bool]
threshold: Optional[float]
strict_mode: bool = Field(alias="strictMode")
skipped: bool
flaky: bool
evaluation_model: Optional[str] = Field(alias="evaluationModel")
evaluation_cost: Optional[float] = Field(alias="evaluationCost")
error: Optional[str]
error_type: Optional[EvaluationErrorType] = Field(alias="errorType")
created_at: str = Field(alias="createdAt")
evaluated_at: Optional[str] = Field(alias="evaluatedAt")
multi_turn: bool = Field(alias="multiTurn")
trace_uuid: Optional[str] = Field(alias="traceUuid")
span_uuid: Optional[str] = Field(alias="spanUuid")
thread_id: Optional[str] = Field(alias="threadId")
test_case_id: Optional[str] = Field(alias="testCaseId")
test_run_id: Optional[str] = Field(alias="testRunId")idstrRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestrRequired
The name of the metric.
Example: "Answer Relevancy"
scoreOptional[float]Required
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonOptional[str]Required
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successOptional[bool]Required
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdOptional[float]Required
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strict_modeboolRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedboolRequired
Whether the metric evaluation was skipped.
Example: false
flakyboolRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluation_modelOptional[str]Required
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluation_costOptional[float]Required
The cost of running the evaluation in USD.
Example: 0.0004
errorOptional[str]Required
The error message if the evaluation failed.
error_typeOptional[EvaluationErrorType]Required
See EvaluationErrorType.
created_atstrRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluated_atOptional[str]Required
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
multi_turnboolRequired
Whether this metric was evaluated on a multi-turn conversation.
Example: false
trace_uuidOptional[str]Required
The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
span_uuidOptional[str]Required
The uuid of the span this metric was evaluated on, for component-level metrics.
thread_idOptional[str]Required
The id of the thread this metric was evaluated on, for conversation-level metrics.
test_case_idOptional[str]Required
The id of the test case this metric was evaluated on.
test_run_idOptional[str]Required
The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.
interface MetricData {
id: string;
name: string;
score: number | null;
reason: string | null;
success: boolean | null;
threshold: number | null;
strictMode: boolean;
skipped: boolean;
flaky: boolean;
evaluationModel: string | null;
evaluationCost: number | null;
error: string | null;
errorType: EvaluationErrorType | null;
createdAt: string;
evaluatedAt: string | null;
multiTurn: boolean;
traceUuid: string | null;
spanUuid: string | null;
threadId: string | null;
testCaseId: string | null;
testRunId: string | null;
}idstringRequired
The unique identifier of the metric data entry.
Example: "<METRIC-DATA-ID>"
namestringRequired
The name of the metric.
Example: "Answer Relevancy"
scorenumber | nullRequired
The final metric score, or null when the metric errored or was skipped.
Example: 0.95
reasonstring | nullRequired
The reason for the metric score, generated by the evaluation model at evaluation time.
Example: "The answer directly states the capital of France."
successboolean | nullRequired
Whether the metric score is above the threshold, or null while the evaluation is still running.
Example: true
thresholdnumber | nullRequired
The threshold for the metric, which determines if the metric is passing or failing.
Example: 0.5
strictModebooleanRequired
Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.
Example: false
skippedbooleanRequired
Whether the metric evaluation was skipped.
Example: false
flakybooleanRequired
Whether the metric's verdict was non-deterministic across runs.
Example: false
evaluationModelstring | nullRequired
The evaluation model used to run the evaluation.
Example: "gpt-4o"
evaluationCostnumber | nullRequired
The cost of running the evaluation in USD.
Example: 0.0004
errorstring | nullRequired
The error message if the evaluation failed.
errorTypeEvaluationErrorType | nullRequired
See EvaluationErrorType.
createdAtstringRequired
The time the metric data was created.
Example: "2025-01-15T10:30:06+00:00"
evaluatedAtstring | nullRequired
The time the metric was evaluated, or null while it is still running.
Example: "2025-01-15T10:30:09+00:00"
multiTurnbooleanRequired
Whether this metric was evaluated on a multi-turn conversation.
Example: false
traceUuidstring | nullRequired
The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.
Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"
spanUuidstring | nullRequired
The uuid of the span this metric was evaluated on, for component-level metrics.
threadIdstring | nullRequired
The id of the thread this metric was evaluated on, for conversation-level metrics.
testCaseIdstring | nullRequired
The id of the test case this metric was evaluated on.
testRunIdstring | nullRequired
The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.
Span
A span with its full input and output, evaluation fields, results and annotations.
class Span:
uuid: str
trace_uuid: str = Field(alias="traceUuid")
parent_uuid: Optional[str] = Field(alias="parentUuid")
name: Optional[str]
type: SpanType
status: TraceSpanStatus
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
error: Optional[str]
integration: Optional[str]
provider: Optional[str]
model: Optional[str]
endpoint: Optional[str]
cost: Optional[float]
input_token_cost: Optional[float] = Field(alias="inputTokenCost")
output_token_cost: Optional[float] = Field(alias="outputTokenCost")
cost_per_input_token: Optional[float] = Field(alias="costPerInputToken")
cost_per_output_token: Optional[float] = Field(alias="costPerOutputToken")
input_token_count: Optional[int] = Field(alias="inputTokenCount")
output_token_count: Optional[int] = Field(alias="outputTokenCount")
prompt_alias: Optional[str] = Field(alias="promptAlias")
prompt_version: Optional[str] = Field(alias="promptVersion")
prompt_label: Optional[str] = Field(alias="promptLabel")
prompt_commit_hash: Optional[str] = Field(alias="promptCommitHash")
embedder: Optional[str]
top_k: Optional[int] = Field(alias="topK")
chunk_size: Optional[int] = Field(alias="chunkSize")
description: Optional[str]
agent_handoffs: Optional[List[str]] = Field(alias="agentHandoffs")
available_tools: Optional[List[str]] = Field(alias="availableTools")
metadata: Optional[Dict[str, Any]]
metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
input: Optional[str]
output: Optional[str]
expected_output: Optional[str] = Field(alias="expectedOutput")
retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
context: Optional[List[str]]
tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
metrics_data: List[MetricData] = Field(alias="metricsData")
annotations: List[AnnotationSummary]uuidstrRequired
This is the unique identifier of the span.
Example: "<SPAN-UUID>"
trace_uuidstrRequired
This is the uuid of the trace containing the span.
Example: "<TRACE-UUID>"
parent_uuidOptional[str]Required
This is the uuid of the parent span, or null for a root span.
Example: "<PARENT-SPAN-UUID>"
nameOptional[str]Required
This is the name of the span.
Example: "OpenAI Call"
typeSpanTypeRequired
See SpanType.
statusTraceSpanStatusRequired
See TraceSpanStatus.
start_timestrRequired
This is the time the span started.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the span ended.
Example: "2025-01-15T10:30:02+00:00"
errorOptional[str]Required
This is the error string that caused the span to fail, or null when no error occurred.
integrationOptional[str]Required
This is the integration associated with the span.
Example: "LangChain"
providerOptional[str]Required
This is the LLM provider used in an LLM span.
Example: "OpenAI"
modelOptional[str]Required
This is the LLM model used in an LLM span.
Example: "gpt-4o"
endpointOptional[str]Required
This is the API endpoint the model was called through in an LLM span.
costOptional[float]Required
This is the total cost of the span in USD, or null when it is not known.
Example: 0.00018
input_token_costOptional[float]Required
This is the total cost of the input tokens passed to the LLM model in an LLM span.
Example: 0.00006
output_token_costOptional[float]Required
This is the total cost of the output tokens generated by the LLM model in an LLM span.
Example: 0.00012
cost_per_input_tokenOptional[float]Required
This is the cost per input token of the LLM model for an LLM span.
Example: 0.0000025
cost_per_output_tokenOptional[float]Required
This is the cost per output token of the LLM model for an LLM span.
Example: 0.00001
input_token_countOptional[int]Required
This is the total number of input tokens passed to the LLM model in an LLM span.
Example: 24
output_token_countOptional[int]Required
This is the total number of output tokens generated by the LLM model in an LLM span.
Example: 12
prompt_aliasOptional[str]Required
This is the alias of your prompt which is stored on Confident AI.
Example: "geography-assistant"
prompt_versionOptional[str]Required
This is the version assigned to your prompt on Confident AI.
Example: "00.00.01"
prompt_labelOptional[str]Required
This is the label assigned to a specific version of prompt on the Confident AI platform.
Example: "production"
prompt_commit_hashOptional[str]Required
This is the hash of the current prompt being logged in the llm span.
Example: "bab04ce"
embedderOptional[str]Required
This is the embedder model used in a retriever span.
top_kOptional[int]Required
This is the top K chunks retrieved from your knowledge base in a retriever span.
chunk_sizeOptional[int]Required
This is the chunk size of each retrieved context for a retriever span.
descriptionOptional[str]Required
This is a description if the span is a tool span.
agent_handoffsOptional[List[str]]Required
This is the list of agent handoffs associated with an agent span.
available_toolsOptional[List[str]]Required
This is the list of available tools associated with an agent span.
metadataOptional[Dict[str, Any]]Required
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metric_collection_nameOptional[str]Required
This is the name of the metric collection to evaluate the span.
Example: "LLM Collection Name"
inputOptional[str]Required
This is the input to the span. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputOptional[str]Required
This is the output of the span. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expected_outputOptional[str]Required
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
retrieval_contextOptional[List[str]]Required
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]Required
This is the ideal retrieval context of your span, which is to be used for evaluation.
tools_calledOptional[List[ToolCall]]Required
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expected_toolsOptional[List[ToolCall]]Required
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
metrics_dataList[MetricData]Required
This is the metrics data associated with the span.
See MetricData.
annotationsList[AnnotationSummary]Required
This is the list of annotations associated with the span.
See AnnotationSummary.
interface Span {
uuid: string;
traceUuid: string;
parentUuid: string | null;
name: string | null;
type: SpanType;
status: TraceSpanStatus;
startTime: string;
endTime: string;
error: string | null;
integration: string | null;
provider: string | null;
model: string | null;
endpoint: string | null;
cost: number | null;
inputTokenCost: number | null;
outputTokenCost: number | null;
costPerInputToken: number | null;
costPerOutputToken: number | null;
inputTokenCount: number | null;
outputTokenCount: number | null;
promptAlias: string | null;
promptVersion: string | null;
promptLabel: string | null;
promptCommitHash: string | null;
embedder: string | null;
topK: number | null;
chunkSize: number | null;
description: string | null;
agentHandoffs: string[] | null;
availableTools: string[] | null;
metadata: Record<string, unknown> | null;
metricCollectionName: string | null;
input: string | null;
output: string | null;
expectedOutput: string | null;
retrievalContext: string[] | null;
context: string[] | null;
toolsCalled: ToolCall[] | null;
expectedTools: ToolCall[] | null;
metricsData: MetricData[];
annotations: AnnotationSummary[];
}uuidstringRequired
This is the unique identifier of the span.
Example: "<SPAN-UUID>"
traceUuidstringRequired
This is the uuid of the trace containing the span.
Example: "<TRACE-UUID>"
parentUuidstring | nullRequired
This is the uuid of the parent span, or null for a root span.
Example: "<PARENT-SPAN-UUID>"
namestring | nullRequired
This is the name of the span.
Example: "OpenAI Call"
typeSpanTypeRequired
See SpanType.
statusTraceSpanStatusRequired
See TraceSpanStatus.
startTimestringRequired
This is the time the span started.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the span ended.
Example: "2025-01-15T10:30:02+00:00"
errorstring | nullRequired
This is the error string that caused the span to fail, or null when no error occurred.
integrationstring | nullRequired
This is the integration associated with the span.
Example: "LangChain"
providerstring | nullRequired
This is the LLM provider used in an LLM span.
Example: "OpenAI"
modelstring | nullRequired
This is the LLM model used in an LLM span.
Example: "gpt-4o"
endpointstring | nullRequired
This is the API endpoint the model was called through in an LLM span.
costnumber | nullRequired
This is the total cost of the span in USD, or null when it is not known.
Example: 0.00018
inputTokenCostnumber | nullRequired
This is the total cost of the input tokens passed to the LLM model in an LLM span.
Example: 0.00006
outputTokenCostnumber | nullRequired
This is the total cost of the output tokens generated by the LLM model in an LLM span.
Example: 0.00012
costPerInputTokennumber | nullRequired
This is the cost per input token of the LLM model for an LLM span.
Example: 0.0000025
costPerOutputTokennumber | nullRequired
This is the cost per output token of the LLM model for an LLM span.
Example: 0.00001
inputTokenCountnumber | nullRequired
This is the total number of input tokens passed to the LLM model in an LLM span.
Example: 24
outputTokenCountnumber | nullRequired
This is the total number of output tokens generated by the LLM model in an LLM span.
Example: 12
promptAliasstring | nullRequired
This is the alias of your prompt which is stored on Confident AI.
Example: "geography-assistant"
promptVersionstring | nullRequired
This is the version assigned to your prompt on Confident AI.
Example: "00.00.01"
promptLabelstring | nullRequired
This is the label assigned to a specific version of prompt on the Confident AI platform.
Example: "production"
promptCommitHashstring | nullRequired
This is the hash of the current prompt being logged in the llm span.
Example: "bab04ce"
embedderstring | nullRequired
This is the embedder model used in a retriever span.
topKnumber | nullRequired
This is the top K chunks retrieved from your knowledge base in a retriever span.
chunkSizenumber | nullRequired
This is the chunk size of each retrieved context for a retriever span.
descriptionstring | nullRequired
This is a description if the span is a tool span.
agentHandoffsstring[] | nullRequired
This is the list of agent handoffs associated with an agent span.
availableToolsstring[] | nullRequired
This is the list of available tools associated with an agent span.
metadataRecord<string, unknown> | nullRequired
This is any additional metadata associated with the span.
Example: {"region":"Europe"}
metricCollectionNamestring | nullRequired
This is the name of the metric collection to evaluate the span.
Example: "LLM Collection Name"
inputstring | nullRequired
This is the input to the span. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputstring | nullRequired
This is the output of the span. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expectedOutputstring | nullRequired
This is the expected output of your span, which is the ideal actual output and to be used for evaluation.
Example: "Paris"
retrievalContextstring[] | nullRequired
This is the retrieval context of your span, which is to be used for evaluation.
Example: ["Paris is the capital and most populous city of France."]
contextstring[] | nullRequired
This is the ideal retrieval context of your span, which is to be used for evaluation.
toolsCalledToolCall[] | nullRequired
This is the tools called by your span, which is to be used for evaluation.
See ToolCall.
expectedToolsToolCall[] | nullRequired
This is the expected tools to be called by the span, which is to be used for evaluation.
See ToolCall.
metricsDataMetricData[]Required
This is the metrics data associated with the span.
See MetricData.
annotationsAnnotationSummary[]Required
This is the list of annotations associated with the span.
See AnnotationSummary.
SpanType
The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.
class SpanType(Enum):
SPAN = "SPAN"
AGENT = "AGENT"
TOOL = "TOOL"
RETRIEVER = "RETRIEVER"
LLM = "LLM"enum SpanType {
SPAN = "SPAN",
AGENT = "AGENT",
TOOL = "TOOL",
RETRIEVER = "RETRIEVER",
LLM = "LLM",
}SPAN · AGENT · TOOL · RETRIEVER · LLM
Thread
A thread with its evaluation results, annotations and the traces that make up the conversation.
class Thread:
id: str
created_at: str = Field(alias="createdAt")
last_activity: str = Field(alias="lastActivity")
metadata: Optional[Dict[str, Any]]
tags: Optional[List[str]]
labels: Dict[str, Classification]
metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
total_traces: int = Field(alias="totalTraces")
metrics_data: List[MetricData] = Field(alias="metricsData")
annotations: List[AnnotationSummary]
traces: List[Trace]idstrRequired
This is the thread id you supplied when creating the thread.
Example: "thread-42"
created_atstrRequired
This is when the thread was created.
Example: "2025-01-15T10:30:00+00:00"
last_activitystrRequired
This is when the thread was last active.
Example: "2025-01-15T11:45:00+00:00"
metadataOptional[Dict[str, Any]]Required
This is the custom metadata attached to the thread.
Example: {"client":"acme-corp","agentId":"geography-agent"}
tagsOptional[List[str]]Required
This is the list of tags associated with the thread.
Example: ["vip"]
labelsDict[str, Classification]Required
The labels your project's classifiers assigned to the thread, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital cities of several countries."}}
metric_collection_nameOptional[str]Required
This is the name of the metric collection assigned to evaluate the thread.
Example: "Conversation Collection Name"
total_tracesintRequired
This is the total number of traces in this thread.
Example: 2
metrics_dataList[MetricData]Required
This is the evaluation metrics data for the thread.
See MetricData.
annotationsList[AnnotationSummary]Required
This is the list of annotations associated with the thread.
See AnnotationSummary.
tracesList[Trace]Required
This is the list of traces in this thread, oldest first and capped at the first 100. Each trace carries its evaluation results and annotations but not its spans.
See Trace.
interface Thread {
id: string;
createdAt: string;
lastActivity: string;
metadata: Record<string, unknown> | null;
tags: string[] | null;
labels: Record<string, Classification>;
metricCollectionName: string | null;
totalTraces: number;
metricsData: MetricData[];
annotations: AnnotationSummary[];
traces: Trace[];
}idstringRequired
This is the thread id you supplied when creating the thread.
Example: "thread-42"
createdAtstringRequired
This is when the thread was created.
Example: "2025-01-15T10:30:00+00:00"
lastActivitystringRequired
This is when the thread was last active.
Example: "2025-01-15T11:45:00+00:00"
metadataRecord<string, unknown> | nullRequired
This is the custom metadata attached to the thread.
Example: {"client":"acme-corp","agentId":"geography-agent"}
tagsstring[] | nullRequired
This is the list of tags associated with the thread.
Example: ["vip"]
labelsRecord<string, Classification>Required
The labels your project's classifiers assigned to the thread, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital cities of several countries."}}
metricCollectionNamestring | nullRequired
This is the name of the metric collection assigned to evaluate the thread.
Example: "Conversation Collection Name"
totalTracesnumberRequired
This is the total number of traces in this thread.
Example: 2
metricsDataMetricData[]Required
This is the evaluation metrics data for the thread.
See MetricData.
annotationsAnnotationSummary[]Required
This is the list of annotations associated with the thread.
See AnnotationSummary.
tracesTrace[]Required
This is the list of traces in this thread, oldest first and capped at the first 100. Each trace carries its evaluation results and annotations but not its spans.
See Trace.
ThreadList
class ThreadList:
threads: List[ThreadSummary]
total_threads: Optional[int] = Field(default=None, alias="totalThreads")
next_cursor: Optional[str] = Field(alias="nextCursor")threadsList[ThreadSummary]Required
This is the list of threads for the current page.
See ThreadSummary.
total_threadsOptional[int]
This is the total number of threads matching the query across all pages. Present on the first page only; omitted when a cursor is given.
Example: 1
next_cursorOptional[str]Required
The value to pass as cursor to get the next page, or null when this is the last page.
interface ThreadList {
threads: ThreadSummary[];
totalThreads?: number;
nextCursor: string | null;
}threadsThreadSummary[]Required
This is the list of threads for the current page.
See ThreadSummary.
totalThreadsnumber
This is the total number of threads matching the query across all pages. Present on the first page only; omitted when a cursor is given.
Example: 1
nextCursorstring | nullRequired
The value to pass as cursor to get the next page, or null when this is the last page.
ThreadSortBy
The thread field to sort by: lastActivity orders by the time of the thread's most recent trace and createdAt by when the thread was created.
class ThreadSortBy(Enum):
LASTACTIVITY = "lastActivity"
CREATEDAT = "createdAt"enum ThreadSortBy {
LASTACTIVITY = "lastActivity",
CREATEDAT = "createdAt",
}LASTACTIVITY · CREATEDAT
ThreadSummary
A thread as it appears in a list: its activity, metadata, tags and labels, without its traces, evaluation results or annotations.
class ThreadSummary:
id: str
created_at: str = Field(alias="createdAt")
last_activity: str = Field(alias="lastActivity")
metadata: Optional[Dict[str, Any]]
tags: Optional[List[str]]
labels: Dict[str, Classification]
metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
total_traces: int = Field(alias="totalTraces")idstrRequired
This is the thread id you supplied when creating the thread.
Example: "thread-42"
created_atstrRequired
This is when the thread was created.
Example: "2025-01-15T10:30:00+00:00"
last_activitystrRequired
This is when the thread was last active.
Example: "2025-01-15T11:45:00+00:00"
metadataOptional[Dict[str, Any]]Required
This is the custom metadata attached to the thread.
Example: {"client":"acme-corp","agentId":"geography-agent"}
tagsOptional[List[str]]Required
This is the list of tags associated with the thread.
Example: ["vip"]
labelsDict[str, Classification]Required
The labels your project's classifiers assigned to the thread, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital cities of several countries."}}
metric_collection_nameOptional[str]Required
This is the name of the metric collection assigned to evaluate the thread.
Example: "Conversation Collection Name"
total_tracesintRequired
This is the total number of traces in this thread.
Example: 2
interface ThreadSummary {
id: string;
createdAt: string;
lastActivity: string;
metadata: Record<string, unknown> | null;
tags: string[] | null;
labels: Record<string, Classification>;
metricCollectionName: string | null;
totalTraces: number;
}idstringRequired
This is the thread id you supplied when creating the thread.
Example: "thread-42"
createdAtstringRequired
This is when the thread was created.
Example: "2025-01-15T10:30:00+00:00"
lastActivitystringRequired
This is when the thread was last active.
Example: "2025-01-15T11:45:00+00:00"
metadataRecord<string, unknown> | nullRequired
This is the custom metadata attached to the thread.
Example: {"client":"acme-corp","agentId":"geography-agent"}
tagsstring[] | nullRequired
This is the list of tags associated with the thread.
Example: ["vip"]
labelsRecord<string, Classification>Required
The labels your project's classifiers assigned to the thread, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital cities of several countries."}}
metricCollectionNamestring | nullRequired
This is the name of the metric collection assigned to evaluate the thread.
Example: "Conversation Collection Name"
totalTracesnumberRequired
This is the total number of traces in this thread.
Example: 2
ToolCall
A tool your LLM application invoked, with what it passed in and what came back.
class ToolCall:
name: str
type: Optional[ToolCallType] = None
description: Optional[str] = None
input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
output: Optional[Any] = None
reasoning: Optional[str] = NonenamestrRequired
This is the name of the tool.
Example: "get_landmark_info"
typeOptional[ToolCallType]
See ToolCallType.
descriptionOptional[str]
This is the description of the tool.
Example: "This tool gives information about a mountain."
input_parametersOptional[Dict[str, Any]]
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputOptional[Any]
This is the output of the tool.
Example: "8,848 metres"
reasoningOptional[str]
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
interface ToolCall {
name: string;
type?: ToolCallType;
description?: string;
inputParameters?: Record<string, unknown> | null;
output?: unknown;
reasoning?: string;
}namestringRequired
This is the name of the tool.
Example: "get_landmark_info"
typeToolCallType
See ToolCallType.
descriptionstring
This is the description of the tool.
Example: "This tool gives information about a mountain."
inputParametersRecord<string, unknown> | null
This is the input parameters that are passed to the tool.
Example: {"mountain":"Everest"}
outputunknown
This is the output of the tool.
Example: "8,848 metres"
reasoningstring
This is the reasoning your LLM provided for the tool call.
Example: "The user asked for the height of a mountain."
ToolCallType
The type of the tool call, either a function or an MCP tool.
class ToolCallType(Enum):
FUNCTION = "FUNCTION"
MCP = "MCP"enum ToolCallType {
FUNCTION = "FUNCTION",
MCP = "MCP",
}FUNCTION · MCP
Trace
A trace with its full input and output, evaluation fields, classifier labels, results and annotations, and its spans when retrieved by id.
class Trace:
uuid: str
name: Optional[str]
status: TraceSpanStatus
start_time: str = Field(alias="startTime")
end_time: str = Field(alias="endTime")
latency: int
cost: Optional[float]
thread_id: Optional[str] = Field(alias="threadId")
user_id: Optional[str] = Field(alias="userId")
customer_id: Optional[str] = Field(alias="customerId")
environment: Environment
tags: Optional[List[str]]
metadata: Optional[Dict[str, Any]]
input: Optional[str]
output: Optional[str]
expected_output: Optional[str] = Field(alias="expectedOutput")
retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
context: Optional[List[str]]
tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
test_case_id: Optional[str] = Field(alias="testCaseId")
metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
labels: Dict[str, Classification]
spans: Optional[List[Span]] = None
metrics_data: List[MetricData] = Field(alias="metricsData")
annotations: List[AnnotationSummary]uuidstrRequired
This is the unique identifier of the trace.
Example: "<TRACE-UUID>"
nameOptional[str]Required
This is the name of the trace.
Example: "Geography QA"
statusTraceSpanStatusRequired
See TraceSpanStatus.
start_timestrRequired
This is the time the trace started.
Example: "2025-01-15T10:30:00+00:00"
end_timestrRequired
This is the time the trace ended.
Example: "2025-01-15T10:30:05+00:00"
latencyintRequired
This is how long the trace took, in milliseconds.
Example: 5000
costOptional[float]Required
This is the total cost of the trace in USD, summed from its spans, or null when it is not known.
Example: 0.00018
thread_idOptional[str]Required
This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.
Example: "thread-42"
user_idOptional[str]Required
This is the user id you provided for this trace, or null when you did not.
Example: "end-user-42"
customer_idOptional[str]Required
This is the customer id you provided for this trace, or null when you did not.
Example: "acme-hotels"
environmentEnvironmentRequired
See Environment.
tagsOptional[List[str]]Required
This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.
Example: ["geography"]
metadataOptional[Dict[str, Any]]Required
This is any additional metadata associated with the trace.
Example: {"client":"acme-corp"}
inputOptional[str]Required
This is the input to the trace. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputOptional[str]Required
This is the output of the trace. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expected_outputOptional[str]Required
This is the expected output associated with the trace, to be used for evaluations.
Example: "Paris"
retrieval_contextOptional[List[str]]Required
This is the retrieval context associated with the trace, to be used for evaluations.
Example: ["Paris is the capital and most populous city of France."]
contextOptional[List[str]]Required
This is the ideal retrieval context associated with the trace, to be used for evaluations.
tools_calledOptional[List[ToolCall]]Required
This is the list of tools called by the trace, to be used for evaluations.
See ToolCall.
expected_toolsOptional[List[ToolCall]]Required
This is the list of expected tools associated with the trace, to be used for evaluations.
See ToolCall.
test_case_idOptional[str]Required
This is the test case id of the trace, which is only set if the trace was created while evaluating a test case.
metric_collection_nameOptional[str]Required
This is the name of the metric collection assigned to evaluate the trace.
Example: "Collection Name"
labelsDict[str, Classification]Required
The labels your project's classifiers assigned to the trace, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital city of a country."}}
spansOptional[List[Span]]
This is the list of spans in the trace, present when the trace is retrieved by id. A thread's traces omit their spans.
See Span.
metrics_dataList[MetricData]Required
This is the list of metrics data associated with the trace after running evaluations.
See MetricData.
annotationsList[AnnotationSummary]Required
This is the list of annotations associated with the trace.
See AnnotationSummary.
interface Trace {
uuid: string;
name: string | null;
status: TraceSpanStatus;
startTime: string;
endTime: string;
latency: number;
cost: number | null;
threadId: string | null;
userId: string | null;
customerId: string | null;
environment: Environment;
tags: string[] | null;
metadata: Record<string, unknown> | null;
input: string | null;
output: string | null;
expectedOutput: string | null;
retrievalContext: string[] | null;
context: string[] | null;
toolsCalled: ToolCall[] | null;
expectedTools: ToolCall[] | null;
testCaseId: string | null;
metricCollectionName: string | null;
labels: Record<string, Classification>;
spans?: Span[];
metricsData: MetricData[];
annotations: AnnotationSummary[];
}uuidstringRequired
This is the unique identifier of the trace.
Example: "<TRACE-UUID>"
namestring | nullRequired
This is the name of the trace.
Example: "Geography QA"
statusTraceSpanStatusRequired
See TraceSpanStatus.
startTimestringRequired
This is the time the trace started.
Example: "2025-01-15T10:30:00+00:00"
endTimestringRequired
This is the time the trace ended.
Example: "2025-01-15T10:30:05+00:00"
latencynumberRequired
This is how long the trace took, in milliseconds.
Example: 5000
costnumber | nullRequired
This is the total cost of the trace in USD, summed from its spans, or null when it is not known.
Example: 0.00018
threadIdstring | nullRequired
This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.
Example: "thread-42"
userIdstring | nullRequired
This is the user id you provided for this trace, or null when you did not.
Example: "end-user-42"
customerIdstring | nullRequired
This is the customer id you provided for this trace, or null when you did not.
Example: "acme-hotels"
environmentEnvironmentRequired
See Environment.
tagsstring[] | nullRequired
This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.
Example: ["geography"]
metadataRecord<string, unknown> | nullRequired
This is any additional metadata associated with the trace.
Example: {"client":"acme-corp"}
inputstring | nullRequired
This is the input to the trace. JSON inputs are serialized to a string.
Example: "What is the capital of France?"
outputstring | nullRequired
This is the output of the trace. JSON outputs are serialized to a string.
Example: "The capital of France is Paris."
expectedOutputstring | nullRequired
This is the expected output associated with the trace, to be used for evaluations.
Example: "Paris"
retrievalContextstring[] | nullRequired
This is the retrieval context associated with the trace, to be used for evaluations.
Example: ["Paris is the capital and most populous city of France."]
contextstring[] | nullRequired
This is the ideal retrieval context associated with the trace, to be used for evaluations.
toolsCalledToolCall[] | nullRequired
This is the list of tools called by the trace, to be used for evaluations.
See ToolCall.
expectedToolsToolCall[] | nullRequired
This is the list of expected tools associated with the trace, to be used for evaluations.
See ToolCall.
testCaseIdstring | nullRequired
This is the test case id of the trace, which is only set if the trace was created while evaluating a test case.
metricCollectionNamestring | nullRequired
This is the name of the metric collection assigned to evaluate the trace.
Example: "Collection Name"
labelsRecord<string, Classification>Required
The labels your project's classifiers assigned to the trace, keyed by classifier name.
See Classification.
Example: {"intent":{"label":"geography","reason":"The user asks for the capital city of a country."}}
spansSpan[]
This is the list of spans in the trace, present when the trace is retrieved by id. A thread's traces omit their spans.
See Span.
metricsDataMetricData[]Required
This is the list of metrics data associated with the trace after running evaluations.
See MetricData.
annotationsAnnotationSummary[]Required
This is the list of annotations associated with the trace.
See AnnotationSummary.
TraceSpanStatus
This represents the error status of a trace or span: SUCCESS when it completed, ERRORED when it failed.
class TraceSpanStatus(Enum):
SUCCESS = "SUCCESS"
ERRORED = "ERRORED"enum TraceSpanStatus {
SUCCESS = "SUCCESS",
ERRORED = "ERRORED",
}SUCCESS · ERRORED
UserReference
A Confident AI user, as referenced by the records they created.
class UserReference:
id: str
email: str
name: Optional[str]
image: Optional[str]idstrRequired
This is the id of the user.
Example: "<USER-ID>"
emailstrRequired
This is the email address of the user.
Example: "jane@acme.com"
nameOptional[str]Required
This is the display name of the user, or null when they have not set one.
Example: "Jane Doe"
imageOptional[str]Required
This is the URL of the user's avatar, or null when they have none.
interface UserReference {
id: string;
email: string;
name: string | null;
image: string | null;
}idstringRequired
This is the id of the user.
Example: "<USER-ID>"
emailstringRequired
This is the email address of the user.
Example: "jane@acme.com"
namestring | nullRequired
This is the display name of the user, or null when they have not set one.
Example: "Jane Doe"
imagestring | nullRequired
This is the URL of the user's avatar, or null when they have none.
Last updated on