Launch Week 3: Five days of launches

Traces

Overview

The Confident AI SDK exposes every Trace method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Traces

Lists the traces in your Confident AI project one page at a time, newest first by default, as summary rows. Retrieve a trace by uuid for its full detail.

from confident_ai import ConfidentAI
from confident_ai.common import Environment
from confident_ai.traces import TraceSortBy

client = ConfidentAI()

result = client.traces.list(
    page_size=25,
    cursor="<NEXT-CURSOR>",
    start="2025-01-01T00:00:00+00:00",
    end="2025-01-31T23:59:59+00:00",
    ascending="false",
    sort_by=TraceSortBy.CREATEDAT,
    environment=Environment.PRODUCTION,
    metadata={"client": "acme-corp"},
)

For async mode, call a_list and await it as shown below:

result = await client.traces.a_list(...)

Parameters

ParameterTypeDescription
page_sizeOptional[int]The number of results per page, at most 100. Defaults to 25.
cursorOptional[str]This is used for pagination, and should be set to the nextCursor value returned in the previous response to get the next page of results.
startOptional[str]This filters for results created at or after the specified start datetime, in ISO 8601 format. Defaults to 60 days ago.
endOptional[str]This filters for results created before the specified end datetime, in ISO 8601 format. Defaults to the current time.
ascendingOptional[Literal['true', 'false']]This determines if the field specified in sortBy should be in ascending order. Defaults to false, which returns the newest results first.
sort_byOptional[TraceSortBy]This determines the field to sort by. Defaults to createdAt. See TraceSortBy.
environmentOptional[Environment]This filters the traces by the environment where the trace was created, and returns traces from all environments if not specified. See Environment.
metadataOptional[Dict[str, str]]Filter traces by metadata key-value pairs using bracket notation, for example metadata[client]=acme-corp. Every pair must match.

Returns

This method returns an object of type TraceList.

Create Trace

Creates a trace in your Confident AI project, along with the spans it contains, and returns its uuid. A uuid that is not a UUID is hashed into one, so the returned value is what every later lookup uses.

from confident_ai import ConfidentAI
from confident_ai.traces import CustomerRequest
from confident_ai.common import Environment
from confident_ai.common import EvaluationErrorType
from confident_ai.traces import MetricDataConfig
from confident_ai.traces import ThreadRequest
from confident_ai.common import ToolCall
from confident_ai.common import ToolCallType
from confident_ai.common import TraceSpanStatus
from confident_ai.traces import UserRequest

client = ConfidentAI()

result = client.traces.create(
    uuid="<TRACE-UUID>",
    start_time="2025-01-15T10:30:00+00:00",
    end_time="2025-01-15T10:30:05+00:00",
    name="Geography QA",
    input="What is the capital of France?",
    output="The capital of France is Paris.",
    status=TraceSpanStatus.SUCCESS,
    environment=Environment.PRODUCTION,
    metadata={"client": "acme-corp"},
    tags=["geography"],
    thread_id="thread-42",
    thread=ThreadRequest(
        id="thread-42",
        metadata={"client": "acme-corp", "agentId": "geography-agent"},
        tags=["vip"]
    ),
    user_id="end-user-42",
    user=UserRequest(
        id="end-user-42",
        name="Marta Ruiz"
    ),
    customer_id="acme-hotels",
    customer=CustomerRequest(
        id="acme-hotels",
        name="Acme Hotels"
    ),
    metric_collection="Collection Name",
    test_run_id="<TEST-RUN-ID>",
    test_case_id="<TEST-CASE-ID>",
    turn_id="<TURN-ID>",
    retrieval_context=[
        "Paris is the capital and most populous city of France."
    ],
    context=["Paris is the capital of France."],
    expected_output="Paris",
    tools_called=[
        ToolCall(
            name="get_landmark_info",
            type=ToolCallType.FUNCTION,
            description="This tool gives information about a mountain.",
            input_parameters={"mountain": "Everest"},
            output="8,848 metres",
            reasoning="The user asked for the height of a mountain."
        )
    ],
    expected_tools=[
        ToolCall(
            name="get_landmark_info",
            type=ToolCallType.FUNCTION,
            description="This tool gives information about a mountain.",
            input_parameters={"mountain": "Everest"},
            output="8,848 metres",
            reasoning="The user asked for the height of a mountain."
        )
    ],
    spans=[
        {
            "uuid": "<SPAN-UUID>",
            "type": "LLM",
            "name": "OpenAI Call",
            "model": "gpt-4o",
            "provider": "OpenAI",
            "integration": "LangChain",
            "input": "What is the capital of France?",
            "output": "The capital of France is Paris.",
            "startTime": "2025-01-15T10:30:00+00:00",
            "endTime": "2025-01-15T10:30:02+00:00"
        }
    ],
    metrics_data=[
        MetricDataConfig(
            name="Answer Relevancy",
            score=0.95,
            success=True,
            threshold=0.5,
            strict_mode=False,
            flaky=False,
            reason="The answer directly states the capital of France.",
            evaluation_model="gpt-4o",
            evaluation_cost=0.0004,
            error="<ERROR>",
            error_type=EvaluationErrorType.AI_CONNECTION_ERROR,
            verbose_logs="<VERBOSE-LOGS>"
        )
    ],
    attachments={
        "doc-1": {"mimeType": "application/pdf", "dataBase64": "JVBERi0xLjQK"}
    },
)

For async mode, call a_create and await it as shown below:

result = await client.traces.a_create(...)

Parameters

ParameterTypeDescription
uuidstrRequired. The unique identifier of the trace, generated by your application. Values that are not UUIDs are hashed into one, and the hashed uuid is what the response and every later lookup use.
start_timestrRequired. This is the time the trace started, as an ISO 8601 datetime.
end_timestrRequired. This is the time the trace ended, as an ISO 8601 datetime.
nameOptional[str]This is the name of the trace.
inputOptional[Any]This is the input to the trace, as a string or any JSON value.
outputOptional[Any]This is the output of the trace, as a string or any JSON value.
statusOptional[TraceSpanStatus]See TraceSpanStatus.
environmentOptional[Environment]See Environment.
metadataOptional[Dict[str, Any]]This is any additional metadata associated with the trace.
tagsOptional[List[str]]This is any tags associated with the trace, which helps with grouping traces and filtering them on the Confident AI platform.
thread_idOptional[str]This is the unique identifier of the thread associated with the trace, which groups traces in the same thread into a conversation.
threadOptional[ThreadRequest]See ThreadRequest.
user_idOptional[str]This is the unique identifier for your end user for the trace.
userOptional[UserRequest]See UserRequest.
customer_idOptional[str]This is the unique identifier of the customer the trace belongs to — the account, tenant or organization your end user belongs to.
customerOptional[CustomerRequest]See CustomerRequest.
metric_collectionOptional[str]This is the metric collection you wish to use to evaluate the trace.
test_run_idOptional[str]This is the unique identifier of the test run to associate the trace with. When set, the trace becomes one test case in that test run and metricCollection is required. It cannot be combined with testCaseId.
test_case_idOptional[str]The id of an existing test case to attach the trace to, when the trace was produced while evaluating that test case.
turn_idOptional[str]The id of the conversational test case turn to attach the trace to, when the trace was produced while evaluating that turn.
retrieval_contextOptional[List[str]]This is the retrieval context of your trace, which is to be used for evaluation.
contextOptional[List[str]]This is the ideal retrieval context of your trace, which is to be used for evaluation.
expected_outputOptional[str]This is the expected output of your trace, which is the ideal actual output and to be used for evaluation.
tools_calledOptional[List[ToolCall]]This is the tools called by your trace, which is to be used for evaluation. See ToolCall.
expected_toolsOptional[List[ToolCall]]This is the expected tools to be called by the trace, which is to be used for evaluation. See ToolCall.
spansOptional[List[SpanRequest]]This is the list of spans in the trace. Each span's type decides which fields it accepts. See SpanRequest.
metrics_dataOptional[List[MetricDataConfig]]Metric results you already computed for this trace, recorded as-is instead of being evaluated by Confident AI. See MetricDataConfig.
attachmentsOptional[Dict[str, TraceAttachment]]Map of attachment ids to payloads for all [DEEPEVAL:IMAGE:…] and [DEEPEVAL:PDF:…] markers in this trace. Define attachments at the trace level with the same ids for the same instances. See TraceAttachment.

Returns

This method returns an object of type TraceRef.

Get Trace

Retrieves a trace by uuid from your Confident AI project, with its spans, full input and output, evaluation fields, classifier labels, results and annotations.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.traces.get(trace_uuid="<TRACE-UUID>")

For async mode, call a_get and await it as shown below:

result = await client.traces.a_get(...)

Parameters

ParameterTypeDescription
trace_uuidstrRequired. The unique identifier of the trace.

Returns

This method returns an object of type Trace.

Types

AnnotationFieldType

The kind of value an annotation holds: TEXT, NUMBER, FLOAT or BOOLEAN for a free value, CHOICE or MULTIPLE_CHOICE for a choice from the options in the field's config, and THUMBS_RATING or FIVE_STAR_RATING for a rating.

class AnnotationFieldType(Enum):
    TEXT = "TEXT"
    NUMBER = "NUMBER"
    FLOAT = "FLOAT"
    BOOLEAN = "BOOLEAN"
    CHOICE = "CHOICE"
    MULTIPLE_CHOICE = "MULTIPLE_CHOICE"
    FIVE_STAR_RATING = "FIVE_STAR_RATING"
    THUMBS_RATING = "THUMBS_RATING"

TEXT · NUMBER · FLOAT · BOOLEAN · CHOICE · MULTIPLE_CHOICE · FIVE_STAR_RATING · THUMBS_RATING

AnnotationSummary

A human rating as it appears under the trace, span or thread it was left on, without repeating the ids of that target.

class AnnotationSummary:
    id: str
    field_type: AnnotationFieldType = Field(alias="fieldType")
    value: Optional[Union[str, float, bool, List[str]]]
    name: Optional[str]
    explanation: Optional[str]
    expected_outcome: Optional[str] = Field(alias="expectedOutcome")
    expected_output: Optional[str] = Field(alias="expectedOutput")
    created_at: str = Field(alias="createdAt")
    user: Optional[UserReference]

idstrRequired

This is the id of the annotation generated by Confident AI.

Example: "<ANNOTATION-ID>"

field_typeAnnotationFieldTypeRequired

valueOptional[Union[str, float, bool, List[str]]]Required

The annotated value, in the shape its fieldType expects: a boolean for THUMBS_RATING (true is thumbs up) and BOOLEAN, a whole number from 1 to 5 for FIVE_STAR_RATING, a number for NUMBER or FLOAT, a string for TEXT and CHOICE, and a list of strings for MULTIPLE_CHOICE. Null when the annotation was left without a value.

Example: true

nameOptional[str]Required

The name of the annotation.

explanationOptional[str]Required

This is the explanation for the annotation.

Example: "Correct and concise."

expected_outcomeOptional[str]Required

This is the annotated expected outcome, for conversation annotations.

expected_outputOptional[str]Required

This is the annotated expected output, for span and trace annotations.

Example: "The capital of France is Paris."

created_atstrRequired

The timestamp when the annotation was created.

Example: "2025-01-15T11:00:00+00:00"

userOptional[UserReference]Required

Classification

A label assigned by one of your project's classifiers, with the reason it was chosen.

class Classification:
    label: str
    reason: str

labelstrRequired

The label the classifier assigned.

Example: "geography"

reasonstrRequired

The classifier's reason for choosing the label.

Example: "The user asks for the capital city of a country."

CustomerRequest

Customer-level fields applied to the customer record. id is an alternate way to specify the customer and must match top-level customerId if both are provided. name only takes effect when a customer id is resolvable; successive ingestions take the latest name.

class CustomerRequest:
    id: Optional[str] = None
    name: Optional[str] = None

idOptional[str]

The customer id. Equivalent to top-level customerId; if both are set they must match.

Example: "acme-hotels"

nameOptional[str]

A human-readable display name for the customer, shown instead of the id.

Example: "Acme Hotels"

Environment

This is the environment where your trace was posted, which helps with separating and debugging traces from different environments on the Confident AI platform.

class Environment(Enum):
    PRODUCTION = "production"
    DEVELOPMENT = "development"
    STAGING = "staging"
    TESTING = "testing"

PRODUCTION · DEVELOPMENT · STAGING · TESTING

EvaluationErrorType

Why an evaluation errored: the AI connection or a transformer failed, the evaluation model failed, the test case lacked the parameters the metric needs, or an internal error occurred.

class EvaluationErrorType(Enum):
    AI_CONNECTION_ERROR = "AI_CONNECTION_ERROR"
    TRANSFORMER_ERROR = "TRANSFORMER_ERROR"
    EVALUATION_MODEL_ERROR = "EVALUATION_MODEL_ERROR"
    INVALID_TEST_CASE_PARAMETERS = "INVALID_TEST_CASE_PARAMETERS"
    INTERNAL_ERROR = "INTERNAL_ERROR"

AI_CONNECTION_ERROR · TRANSFORMER_ERROR · EVALUATION_MODEL_ERROR · INVALID_TEST_CASE_PARAMETERS · INTERNAL_ERROR

MetricData

The result of an evaluated metric, with the ids of whatever it was recorded against.

class MetricData:
    id: str
    name: str
    score: Optional[float]
    reason: Optional[str]
    success: Optional[bool]
    threshold: Optional[float]
    strict_mode: bool = Field(alias="strictMode")
    skipped: bool
    flaky: bool
    evaluation_model: Optional[str] = Field(alias="evaluationModel")
    evaluation_cost: Optional[float] = Field(alias="evaluationCost")
    error: Optional[str]
    error_type: Optional[EvaluationErrorType] = Field(alias="errorType")
    created_at: str = Field(alias="createdAt")
    evaluated_at: Optional[str] = Field(alias="evaluatedAt")
    multi_turn: bool = Field(alias="multiTurn")
    trace_uuid: Optional[str] = Field(alias="traceUuid")
    span_uuid: Optional[str] = Field(alias="spanUuid")
    thread_id: Optional[str] = Field(alias="threadId")
    test_case_id: Optional[str] = Field(alias="testCaseId")
    test_run_id: Optional[str] = Field(alias="testRunId")

idstrRequired

The unique identifier of the metric data entry.

Example: "<METRIC-DATA-ID>"

namestrRequired

The name of the metric.

Example: "Answer Relevancy"

scoreOptional[float]Required

The final metric score, or null when the metric errored or was skipped.

Example: 0.95

reasonOptional[str]Required

The reason for the metric score, generated by the evaluation model at evaluation time.

Example: "The answer directly states the capital of France."

successOptional[bool]Required

Whether the metric score is above the threshold, or null while the evaluation is still running.

Example: true

thresholdOptional[float]Required

The threshold for the metric, which determines if the metric is passing or failing.

Example: 0.5

strict_modeboolRequired

Whether the metric was run in strict mode, which outputs a binary score of 0 or 1.

Example: false

skippedboolRequired

Whether the metric evaluation was skipped.

Example: false

flakyboolRequired

Whether the metric's verdict was non-deterministic across runs.

Example: false

evaluation_modelOptional[str]Required

The evaluation model used to run the evaluation.

Example: "gpt-4o"

evaluation_costOptional[float]Required

The cost of running the evaluation in USD.

Example: 0.0004

errorOptional[str]Required

The error message if the evaluation failed.

error_typeOptional[EvaluationErrorType]Required

created_atstrRequired

The time the metric data was created.

Example: "2025-01-15T10:30:06+00:00"

evaluated_atOptional[str]Required

The time the metric was evaluated, or null while it is still running.

Example: "2025-01-15T10:30:09+00:00"

multi_turnboolRequired

Whether this metric was evaluated on a multi-turn conversation.

Example: false

trace_uuidOptional[str]Required

The uuid of the trace this metric was evaluated on, or null if it was not evaluated on one.

Example: "3f9c2a1e-5b7d-4c8e-9f01-2a3b4c5d6e7f"

span_uuidOptional[str]Required

The uuid of the span this metric was evaluated on, for component-level metrics.

thread_idOptional[str]Required

The id of the thread this metric was evaluated on, for conversation-level metrics.

test_case_idOptional[str]Required

The id of the test case this metric was evaluated on.

test_run_idOptional[str]Required

The id of the test run this result belongs to. Retrieve that test run to read the result alongside the rest of its test cases.

MetricDataConfig

A metric result you computed yourself, recorded on the trace or span as-is instead of being evaluated by Confident AI.

class MetricDataConfig:
    name: str
    score: Optional[float] = None
    success: Optional[bool] = None
    threshold: Optional[float] = None
    strict_mode: Optional[bool] = Field(default=None, alias="strictMode")
    flaky: Optional[bool] = None
    reason: Optional[str] = None
    evaluation_model: Optional[str] = Field(default=None, alias="evaluationModel")
    evaluation_cost: Optional[float] = Field(default=None, alias="evaluationCost")
    error: Optional[str] = None
    error_type: Optional[EvaluationErrorType] = Field(default=None, alias="errorType")
    verbose_logs: Optional[str] = Field(default=None, alias="verboseLogs")

namestrRequired

The name of the metric.

Example: "Answer Relevancy"

scoreOptional[float]

The metric score, typically between 0 and 1.

Example: 0.95

successOptional[bool]

Whether the metric passed its threshold.

Example: true

thresholdOptional[float]

The threshold the metric was scored against.

Example: 0.5

strict_modeOptional[bool]

Whether the metric ran in strict mode, which outputs a binary score of 0 or 1.

Example: false

flakyOptional[bool]

Whether the metric's verdict was non-deterministic across runs.

Example: false

reasonOptional[str]

The reason for the metric score.

Example: "The answer directly states the capital of France."

evaluation_modelOptional[str]

The model used to evaluate the metric.

Example: "gpt-4o"

evaluation_costOptional[float]

The cost of running the evaluation in USD.

Example: 0.0004

errorOptional[str]

The error message if the evaluation failed.

error_typeOptional[EvaluationErrorType]

verbose_logsOptional[str]

Detailed logs from the evaluation.

Span

A span with its full input and output, evaluation fields, results and annotations.

class Span:
    uuid: str
    trace_uuid: str = Field(alias="traceUuid")
    parent_uuid: Optional[str] = Field(alias="parentUuid")
    name: Optional[str]
    type: SpanType
    status: TraceSpanStatus
    start_time: str = Field(alias="startTime")
    end_time: str = Field(alias="endTime")
    error: Optional[str]
    integration: Optional[str]
    provider: Optional[str]
    model: Optional[str]
    endpoint: Optional[str]
    cost: Optional[float]
    input_token_cost: Optional[float] = Field(alias="inputTokenCost")
    output_token_cost: Optional[float] = Field(alias="outputTokenCost")
    cost_per_input_token: Optional[float] = Field(alias="costPerInputToken")
    cost_per_output_token: Optional[float] = Field(alias="costPerOutputToken")
    input_token_count: Optional[int] = Field(alias="inputTokenCount")
    output_token_count: Optional[int] = Field(alias="outputTokenCount")
    prompt_alias: Optional[str] = Field(alias="promptAlias")
    prompt_version: Optional[str] = Field(alias="promptVersion")
    prompt_label: Optional[str] = Field(alias="promptLabel")
    prompt_commit_hash: Optional[str] = Field(alias="promptCommitHash")
    embedder: Optional[str]
    top_k: Optional[int] = Field(alias="topK")
    chunk_size: Optional[int] = Field(alias="chunkSize")
    description: Optional[str]
    agent_handoffs: Optional[List[str]] = Field(alias="agentHandoffs")
    available_tools: Optional[List[str]] = Field(alias="availableTools")
    metadata: Optional[Dict[str, Any]]
    metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
    input: Optional[str]
    output: Optional[str]
    expected_output: Optional[str] = Field(alias="expectedOutput")
    retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
    context: Optional[List[str]]
    tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
    metrics_data: List[MetricData] = Field(alias="metricsData")
    annotations: List[AnnotationSummary]

uuidstrRequired

This is the unique identifier of the span.

Example: "<SPAN-UUID>"

trace_uuidstrRequired

This is the uuid of the trace containing the span.

Example: "<TRACE-UUID>"

parent_uuidOptional[str]Required

This is the uuid of the parent span, or null for a root span.

Example: "<PARENT-SPAN-UUID>"

nameOptional[str]Required

This is the name of the span.

Example: "OpenAI Call"

typeSpanTypeRequired

statusTraceSpanStatusRequired

start_timestrRequired

This is the time the span started.

Example: "2025-01-15T10:30:00+00:00"

end_timestrRequired

This is the time the span ended.

Example: "2025-01-15T10:30:02+00:00"

errorOptional[str]Required

This is the error string that caused the span to fail, or null when no error occurred.

integrationOptional[str]Required

This is the integration associated with the span.

Example: "LangChain"

providerOptional[str]Required

This is the LLM provider used in an LLM span.

Example: "OpenAI"

modelOptional[str]Required

This is the LLM model used in an LLM span.

Example: "gpt-4o"

endpointOptional[str]Required

This is the API endpoint the model was called through in an LLM span.

costOptional[float]Required

This is the total cost of the span in USD, or null when it is not known.

Example: 0.00018

input_token_costOptional[float]Required

This is the total cost of the input tokens passed to the LLM model in an LLM span.

Example: 0.00006

output_token_costOptional[float]Required

This is the total cost of the output tokens generated by the LLM model in an LLM span.

Example: 0.00012

cost_per_input_tokenOptional[float]Required

This is the cost per input token of the LLM model for an LLM span.

Example: 0.0000025

cost_per_output_tokenOptional[float]Required

This is the cost per output token of the LLM model for an LLM span.

Example: 0.00001

input_token_countOptional[int]Required

This is the total number of input tokens passed to the LLM model in an LLM span.

Example: 24

output_token_countOptional[int]Required

This is the total number of output tokens generated by the LLM model in an LLM span.

Example: 12

prompt_aliasOptional[str]Required

This is the alias of your prompt which is stored on Confident AI.

Example: "geography-assistant"

prompt_versionOptional[str]Required

This is the version assigned to your prompt on Confident AI.

Example: "00.00.01"

prompt_labelOptional[str]Required

This is the label assigned to a specific version of prompt on the Confident AI platform.

Example: "production"

prompt_commit_hashOptional[str]Required

This is the hash of the current prompt being logged in the llm span.

Example: "bab04ce"

embedderOptional[str]Required

This is the embedder model used in a retriever span.

top_kOptional[int]Required

This is the top K chunks retrieved from your knowledge base in a retriever span.

chunk_sizeOptional[int]Required

This is the chunk size of each retrieved context for a retriever span.

descriptionOptional[str]Required

This is a description if the span is a tool span.

agent_handoffsOptional[List[str]]Required

This is the list of agent handoffs associated with an agent span.

available_toolsOptional[List[str]]Required

This is the list of available tools associated with an agent span.

metadataOptional[Dict[str, Any]]Required

This is any additional metadata associated with the span.

Example: {"region":"Europe"}

metric_collection_nameOptional[str]Required

This is the name of the metric collection to evaluate the span.

Example: "LLM Collection Name"

inputOptional[str]Required

This is the input to the span. JSON inputs are serialized to a string.

Example: "What is the capital of France?"

outputOptional[str]Required

This is the output of the span. JSON outputs are serialized to a string.

Example: "The capital of France is Paris."

expected_outputOptional[str]Required

This is the expected output of your span, which is the ideal actual output and to be used for evaluation.

Example: "Paris"

retrieval_contextOptional[List[str]]Required

This is the retrieval context of your span, which is to be used for evaluation.

Example: ["Paris is the capital and most populous city of France."]

contextOptional[List[str]]Required

This is the ideal retrieval context of your span, which is to be used for evaluation.

tools_calledOptional[List[ToolCall]]Required

This is the tools called by your span, which is to be used for evaluation.

See ToolCall.

expected_toolsOptional[List[ToolCall]]Required

This is the expected tools to be called by the span, which is to be used for evaluation.

See ToolCall.

metrics_dataList[MetricData]Required

This is the metrics data associated with the span.

See MetricData.

annotationsList[AnnotationSummary]Required

This is the list of annotations associated with the span.

See AnnotationSummary.

SpanRequest

A span to record inside the trace. Its type selects the variant: omit it or send SPAN for a plain span, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, or AGENT for an agent step. Each variant accepts only its own fields.

SpanRequest = Union[
    BaseSpanRequest,
    LlmSpanRequest,
    RetrieverSpanRequest,
    ToolSpanRequest,
    AgentSpanRequest,
]

A SpanRequest is one of the shapes below. Send the fields of one of them, never a mix of both.

A plain span with no model, retriever, tool or agent detail.

class BaseSpanRequest:
    type: Optional[Literal["SPAN"]] = None
    uuid: str
    name: str
    input: Optional[Any] = None
    output: Optional[Any] = None
    error: Optional[str] = None
    status: Optional[TraceSpanStatus] = None
    start_time: str = Field(alias="startTime")
    end_time: str = Field(alias="endTime")
    parent_uuid: Optional[str] = Field(default=None, alias="parentUuid")
    metadata: Optional[Dict[str, Any]] = None
    metric_collection: Optional[str] = Field(default=None, alias="metricCollection")
    retrieval_context: Optional[List[str]] = Field(default=None, alias="retrievalContext")
    context: Optional[List[str]] = None
    expected_output: Optional[str] = Field(default=None, alias="expectedOutput")
    tools_called: Optional[List[ToolCall]] = Field(default=None, alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(default=None, alias="expectedTools")
    integration: Optional[str] = None
    metrics_data: Optional[List[MetricDataConfig]] = Field(default=None, alias="metricsData")

typeOptional[Literal["SPAN"]]

The type of the span. Omit it, or send SPAN, for a plain span.

Example: "SPAN"

uuidstrRequired

The unique identifier of the span, generated by your application. Values that are not UUIDs are hashed into one.

Example: "<SPAN-UUID>"

namestrRequired

This is the name of the span.

Example: "OpenAI Call"

inputOptional[Any]

This is the input to the span, as a string or any JSON value.

Example: "What is the capital of France?"

outputOptional[Any]

This is the output of the span, as a string or any JSON value.

Example: "The capital of France is Paris."

errorOptional[str]

This is the error message, if an error occurred inside the span.

Example: "The model timed out after 30 seconds."

statusOptional[TraceSpanStatus]

start_timestrRequired

This is the time the span started, as an ISO 8601 datetime.

Example: "2025-01-15T10:30:00+00:00"

end_timestrRequired

This is the time the span ended, as an ISO 8601 datetime.

Example: "2025-01-15T10:30:02+00:00"

parent_uuidOptional[str]

This is the unique identifier of the span's parent span. Omit it for a root span.

Example: "<PARENT-SPAN-UUID>"

metadataOptional[Dict[str, Any]]

This is any additional metadata associated with the span.

Example: {"region":"Europe"}

metric_collectionOptional[str]

This is the metric collection to be used for evaluating the span.

Example: "LLM Collection Name"

retrieval_contextOptional[List[str]]

This is the retrieval context of your span, which is to be used for evaluation.

Example: ["Paris is the capital and most populous city of France."]

contextOptional[List[str]]

This is the ideal retrieval context of your span, which is to be used for evaluation.

Example: ["Paris is the capital of France."]

expected_outputOptional[str]

This is the expected output of your span, which is the ideal actual output and to be used for evaluation.

Example: "Paris"

tools_calledOptional[List[ToolCall]]

This is the tools called by your span, which is to be used for evaluation.

See ToolCall.

expected_toolsOptional[List[ToolCall]]

This is the expected tools to be called by the span, which is to be used for evaluation.

See ToolCall.

integrationOptional[str]

This is the integration associated with the span.

Example: "LangChain"

metrics_dataOptional[List[MetricDataConfig]]

Metric results you already computed for this span, recorded as-is instead of being evaluated by Confident AI.

See MetricDataConfig.

SpanType

The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.

class SpanType(Enum):
    SPAN = "SPAN"
    AGENT = "AGENT"
    TOOL = "TOOL"
    RETRIEVER = "RETRIEVER"
    LLM = "LLM"

SPAN · AGENT · TOOL · RETRIEVER · LLM

ThreadRequest

Thread-level fields applied to the thread record. id is an alternate way to specify the thread and must match top-level threadId if both are provided. metadata and tags only take effect when a thread id is resolvable; successive ingestions merge metadata keys, while tags replace any prior value.

class ThreadRequest:
    id: Optional[str] = None
    metadata: Optional[Dict[str, Any]] = None
    tags: Optional[List[str]] = None

idOptional[str]

The thread id. Equivalent to top-level threadId; if both are set they must match.

Example: "thread-42"

metadataOptional[Dict[str, Any]]

Custom key/value metadata to attach to the thread. Values can be any JSON-serializable type and are stringified server-side. Successive ingestions for the same thread merge metadata keys.

Example: {"client":"acme-corp","agentId":"geography-agent"}

tagsOptional[List[str]]

Tags to set on the thread. Replaces any previously stored tags.

Example: ["vip"]

ToolCall

A tool your LLM application invoked, with what it passed in and what came back.

class ToolCall:
    name: str
    type: Optional[ToolCallType] = None
    description: Optional[str] = None
    input_parameters: Optional[Dict[str, Any]] = Field(default=None, alias="inputParameters")
    output: Optional[Any] = None
    reasoning: Optional[str] = None

namestrRequired

This is the name of the tool.

Example: "get_landmark_info"

typeOptional[ToolCallType]

descriptionOptional[str]

This is the description of the tool.

Example: "This tool gives information about a mountain."

input_parametersOptional[Dict[str, Any]]

This is the input parameters that are passed to the tool.

Example: {"mountain":"Everest"}

outputOptional[Any]

This is the output of the tool.

Example: "8,848 metres"

reasoningOptional[str]

This is the reasoning your LLM provided for the tool call.

Example: "The user asked for the height of a mountain."

ToolCallType

The type of the tool call, either a function or an MCP tool.

class ToolCallType(Enum):
    FUNCTION = "FUNCTION"
    MCP = "MCP"

FUNCTION · MCP

Trace

A trace with its full input and output, evaluation fields, classifier labels, results and annotations, and its spans when retrieved by id.

class Trace:
    uuid: str
    name: Optional[str]
    status: TraceSpanStatus
    start_time: str = Field(alias="startTime")
    end_time: str = Field(alias="endTime")
    latency: int
    cost: Optional[float]
    thread_id: Optional[str] = Field(alias="threadId")
    user_id: Optional[str] = Field(alias="userId")
    customer_id: Optional[str] = Field(alias="customerId")
    environment: Environment
    tags: Optional[List[str]]
    metadata: Optional[Dict[str, Any]]
    input: Optional[str]
    output: Optional[str]
    expected_output: Optional[str] = Field(alias="expectedOutput")
    retrieval_context: Optional[List[str]] = Field(alias="retrievalContext")
    context: Optional[List[str]]
    tools_called: Optional[List[ToolCall]] = Field(alias="toolsCalled")
    expected_tools: Optional[List[ToolCall]] = Field(alias="expectedTools")
    test_case_id: Optional[str] = Field(alias="testCaseId")
    metric_collection_name: Optional[str] = Field(alias="metricCollectionName")
    labels: Dict[str, Classification]
    spans: Optional[List[Span]] = None
    metrics_data: List[MetricData] = Field(alias="metricsData")
    annotations: List[AnnotationSummary]

uuidstrRequired

This is the unique identifier of the trace.

Example: "<TRACE-UUID>"

nameOptional[str]Required

This is the name of the trace.

Example: "Geography QA"

statusTraceSpanStatusRequired

start_timestrRequired

This is the time the trace started.

Example: "2025-01-15T10:30:00+00:00"

end_timestrRequired

This is the time the trace ended.

Example: "2025-01-15T10:30:05+00:00"

latencyintRequired

This is how long the trace took, in milliseconds.

Example: 5000

costOptional[float]Required

This is the total cost of the trace in USD, summed from its spans, or null when it is not known.

Example: 0.00018

thread_idOptional[str]Required

This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.

Example: "thread-42"

user_idOptional[str]Required

This is the user id you provided for this trace, or null when you did not.

Example: "end-user-42"

customer_idOptional[str]Required

This is the customer id you provided for this trace, or null when you did not.

Example: "acme-hotels"

environmentEnvironmentRequired

tagsOptional[List[str]]Required

This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.

Example: ["geography"]

metadataOptional[Dict[str, Any]]Required

This is any additional metadata associated with the trace.

Example: {"client":"acme-corp"}

inputOptional[str]Required

This is the input to the trace. JSON inputs are serialized to a string.

Example: "What is the capital of France?"

outputOptional[str]Required

This is the output of the trace. JSON outputs are serialized to a string.

Example: "The capital of France is Paris."

expected_outputOptional[str]Required

This is the expected output associated with the trace, to be used for evaluations.

Example: "Paris"

retrieval_contextOptional[List[str]]Required

This is the retrieval context associated with the trace, to be used for evaluations.

Example: ["Paris is the capital and most populous city of France."]

contextOptional[List[str]]Required

This is the ideal retrieval context associated with the trace, to be used for evaluations.

tools_calledOptional[List[ToolCall]]Required

This is the list of tools called by the trace, to be used for evaluations.

See ToolCall.

expected_toolsOptional[List[ToolCall]]Required

This is the list of expected tools associated with the trace, to be used for evaluations.

See ToolCall.

test_case_idOptional[str]Required

This is the test case id of the trace, which is only set if the trace was created while evaluating a test case.

metric_collection_nameOptional[str]Required

This is the name of the metric collection assigned to evaluate the trace.

Example: "Collection Name"

labelsDict[str, Classification]Required

The labels your project's classifiers assigned to the trace, keyed by classifier name.

See Classification.

Example: {"intent":{"label":"geography","reason":"The user asks for the capital city of a country."}}

spansOptional[List[Span]]

This is the list of spans in the trace, present when the trace is retrieved by id. A thread's traces omit their spans.

See Span.

metrics_dataList[MetricData]Required

This is the list of metrics data associated with the trace after running evaluations.

See MetricData.

annotationsList[AnnotationSummary]Required

This is the list of annotations associated with the trace.

See AnnotationSummary.

TraceAttachment

Payload for a multimodal attachment referenced by id in [DEEPEVAL:IMAGE:…] or [DEEPEVAL:PDF:…] markers. Provide either a url, or dataBase64 with mimeType.

class TraceAttachment:
    url: Optional[str] = None
    data_base64: Optional[str] = Field(default=None, alias="dataBase64")
    mime_type: Optional[str] = Field(default=None, alias="mimeType")

urlOptional[str]

Public URL of the attachment. Send either url, or dataBase64 with mimeType, not both.

Example: "https://example.com/paris.pdf"

data_base64Optional[str]

Base64-encoded file bytes, as an alternative to url.

Example: "JVBERi0xLjQK"

mime_typeOptional[str]

MIME type of the attachment, required when using dataBase64.

Example: "application/pdf"

TraceList

class TraceList:
    traces: List[TraceSummary]
    total_traces: Optional[int] = Field(default=None, alias="totalTraces")
    next_cursor: Optional[str] = Field(alias="nextCursor")

tracesList[TraceSummary]Required

This is the list of traces for the current page.

See TraceSummary.

total_tracesOptional[int]

This is the total number of traces matching the query across all pages. Present on the first page only; omitted when a cursor is given.

Example: 1

next_cursorOptional[str]Required

The value to pass as cursor to get the next page, or null when this is the last page.

TraceRef

class TraceRef:
    uuid: str

uuidstrRequired

This is the uuid of the trace. It is the uuid you sent, or its UUID hash when the value you sent was not a UUID.

Example: "<TRACE-UUID>"

TraceSortBy

The trace field to sort by: createdAt orders by the trace's start time and endedAt by its end time.

class TraceSortBy(Enum):
    CREATEDAT = "createdAt"
    ENDEDAT = "endedAt"

CREATEDAT · ENDEDAT

TraceSpanStatus

This represents the error status of a trace or span: SUCCESS when it completed, ERRORED when it failed.

class TraceSpanStatus(Enum):
    SUCCESS = "SUCCESS"
    ERRORED = "ERRORED"

SUCCESS · ERRORED

TraceSummary

A trace as it appears in a list: its timing, cost, thread and user with a preview of its input and output, but without its spans, evaluation fields, results or annotations.

class TraceSummary:
    uuid: str
    name: Optional[str]
    status: TraceSpanStatus
    start_time: str = Field(alias="startTime")
    end_time: str = Field(alias="endTime")
    latency: int
    cost: Optional[float]
    thread_id: Optional[str] = Field(alias="threadId")
    user_id: Optional[str] = Field(alias="userId")
    customer_id: Optional[str] = Field(alias="customerId")
    environment: Environment
    tags: Optional[List[str]]
    metadata: Optional[Dict[str, Any]]
    input_preview: Optional[str] = Field(alias="inputPreview")
    output_preview: Optional[str] = Field(alias="outputPreview")

uuidstrRequired

This is the unique identifier of the trace.

Example: "<TRACE-UUID>"

nameOptional[str]Required

This is the name of the trace.

Example: "Geography QA"

statusTraceSpanStatusRequired

start_timestrRequired

This is the time the trace started.

Example: "2025-01-15T10:30:00+00:00"

end_timestrRequired

This is the time the trace ended.

Example: "2025-01-15T10:30:05+00:00"

latencyintRequired

This is how long the trace took, in milliseconds.

Example: 5000

costOptional[float]Required

This is the total cost of the trace in USD, summed from its spans, or null when it is not known.

Example: 0.00018

thread_idOptional[str]Required

This is the thread id of the trace, which groups traces in the same thread into a conversation, or null when the trace is not part of one.

Example: "thread-42"

user_idOptional[str]Required

This is the user id you provided for this trace, or null when you did not.

Example: "end-user-42"

customer_idOptional[str]Required

This is the customer id you provided for this trace, or null when you did not.

Example: "acme-hotels"

environmentEnvironmentRequired

tagsOptional[List[str]]Required

This is the list of tags associated with the trace, which is useful for grouping and filtering for traces.

Example: ["geography"]

metadataOptional[Dict[str, Any]]Required

This is any additional metadata associated with the trace.

Example: {"client":"acme-corp"}

input_previewOptional[str]Required

The first characters of the trace's input, or null when it has none. Retrieve the trace by id for the full value.

Example: "What is the capital of France?"

output_previewOptional[str]Required

The first characters of the trace's output, or null when it has none. Retrieve the trace by id for the full value.

Example: "The capital of France is Paris."

UserReference

A Confident AI user, as referenced by the records they created.

class UserReference:
    id: str
    email: str
    name: Optional[str]
    image: Optional[str]

idstrRequired

This is the id of the user.

Example: "<USER-ID>"

emailstrRequired

This is the email address of the user.

Example: "jane@acme.com"

nameOptional[str]Required

This is the display name of the user, or null when they have not set one.

Example: "Jane Doe"

imageOptional[str]Required

This is the URL of the user's avatar, or null when they have none.

UserRequest

End-user-level fields applied to the end-user record. id is an alternate way to specify the end user and must match top-level userId if both are provided. name only takes effect when an end-user id is resolvable; successive ingestions take the latest name.

class UserRequest:
    id: Optional[str] = None
    name: Optional[str] = None

idOptional[str]

The end-user id. Equivalent to top-level userId; if both are set they must match.

Example: "end-user-42"

nameOptional[str]

A human-readable display name for the end user, shown instead of the id.

Example: "Marta Ruiz"

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI