Launch Week 3: Five days of launches

Evaluation Rules

Overview

The Confident AI SDK exposes every Evaluation Rule method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Evaluation Rules

Lists the evaluation rules in your Confident AI project one page at a time, newest first, as summary rows. Requires the Starter plan or above.

from confident_ai import ConfidentAI
from confident_ai.evaluation_rules import EvaluationRuleDataModel

client = ConfidentAI()

result = client.evaluation_rules.list(
    data_model=EvaluationRuleDataModel.TRACE,
    page=1,
    page_size=25,
)

For async mode, call a_list and await it as shown below:

result = await client.evaluation_rules.a_list(...)

Parameters

ParameterTypeDescription
data_modelOptional[EvaluationRuleDataModel]See EvaluationRuleDataModel.
pageOptional[int]The page to return. Defaults to 1.
page_sizeOptional[int]The number of results per page, at most 100. Defaults to 25.

Returns

This method returns an object of type EvaluationRuleList.

Create Evaluation Rule

Creates a standing rule that runs a metric collection against matching production traces, spans or threads as they arrive, and returns its id. Running metrics consumes LLM usage. The collection's turn type must match the rule — THREAD rules need a multi-turn collection, TRACE and SPAN rules a single-turn one — and only one enabled THREAD rule may target a given collection. Requires the Starter plan or above.

from confident_ai import ConfidentAI
from confident_ai.evaluation_rules import EvaluationRuleDataModel
from confident_ai.common import SpanType

client = ConfidentAI()

result = client.evaluation_rules.create(
    name="Score production answers",
    data_model=EvaluationRuleDataModel.TRACE,
    metric_collection_id="<METRIC-COLLECTION-ID>",
    enabled=True,
    description="Scores answers we serve to end users.",
    sample_rate=0.2,
    span_type=SpanType.SPAN,
    filters={
        "operator": "AND",
        "groups": [
            {
                "operator": "AND",
                "filters": [
                    {
                        "category": "Name",
                        "condition": "Is",
                        "value": "capital-lookup"
                    }
                ]
            }
        ]
    },
    thread_timelimit=300,
    overwrite_evals=False,
)

For async mode, call a_create and await it as shown below:

result = await client.evaluation_rules.a_create(...)

Parameters

ParameterTypeDescription
namestrRequired. A name for the rule, unique within the project.
data_modelEvaluationRuleDataModelRequired. See EvaluationRuleDataModel.
metric_collection_idstrRequired. The id of the metric collection to run. It must be multi-turn for THREAD rules and single-turn for TRACE and SPAN rules.
enabledOptional[bool]Whether the rule evaluates matching items as they arrive. Defaults to true.
descriptionOptional[str]A note about what the rule checks. Send null to clear it.
sample_rateOptional[float]The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them.
span_typeOptional[SpanType]Only evaluate spans of this kind. Allowed only when dataModel is SPAN, and cleared automatically if the rule moves off SPAN. Send null to evaluate every span. See SpanType.
filtersOptional[FilterSet]Only evaluate items matching these filters. Send null to evaluate every item the rule's dataModel covers. See FilterSet.
thread_timelimitOptional[int]For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300.
overwrite_evalsOptional[bool]Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false.

Returns

This method returns an object of type EvaluationRuleRef.

Get Evaluation Rule

Retrieves an evaluation rule by id with its full configuration, including the filters an item must match and the id of the metric collection it runs.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.evaluation_rules.get(
    evaluation_rule_id="<EVALUATION-RULE-ID>",
)

For async mode, call a_get and await it as shown below:

result = await client.evaluation_rules.a_get(...)

Parameters

ParameterTypeDescription
evaluation_rule_idstrRequired. The id of the evaluation rule.

Returns

This method returns an object of type EvaluationRule.

Update Evaluation Rule

Updates an evaluation rule and returns it. Only the fields you send are changed; omitting a field leaves it untouched, and sending null clears it. Constraints are re-checked against the rule the update produces, not just the fields you sent, so switching a rule to THREAD still requires a multi-turn metric collection.

from confident_ai import ConfidentAI
from confident_ai.evaluation_rules import EvaluationRuleDataModel
from confident_ai.common import SpanType

client = ConfidentAI()

result = client.evaluation_rules.update(
    evaluation_rule_id="<EVALUATION-RULE-ID>",
    name="Score production answers",
    enabled=False,
    data_model=EvaluationRuleDataModel.TRACE,
    metric_collection_id="<METRIC-COLLECTION-ID>",
    description="Scores answers we serve to end users.",
    sample_rate=0.2,
    span_type=SpanType.SPAN,
    filters={
        "operator": "AND",
        "groups": [
            {
                "operator": "AND",
                "filters": [
                    {
                        "category": "Name",
                        "condition": "Is",
                        "value": "capital-lookup"
                    }
                ]
            }
        ]
    },
    thread_timelimit=300,
    overwrite_evals=False,
)

For async mode, call a_update and await it as shown below:

result = await client.evaluation_rules.a_update(...)

Parameters

ParameterTypeDescription
evaluation_rule_idstrRequired. The id of the evaluation rule.
nameOptional[str]A new name for the rule, unique within the project.
enabledOptional[bool]Whether the rule evaluates matching items as they arrive.
data_modelOptional[EvaluationRuleDataModel]See EvaluationRuleDataModel.
metric_collection_idOptional[str]The id of a different metric collection to run.
descriptionOptional[str]A note about what the rule checks. Send null to clear it.
sample_rateOptional[float]The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them.
span_typeOptional[SpanType]Only evaluate spans of this kind. Allowed only when dataModel is SPAN, and cleared automatically if the rule moves off SPAN. Send null to evaluate every span. See SpanType.
filtersOptional[FilterSet]Only evaluate items matching these filters. Send null to evaluate every item the rule's dataModel covers. See FilterSet.
thread_timelimitOptional[int]For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300.
overwrite_evalsOptional[bool]Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false.

Returns

This method returns an object of type EvaluationRule.

Delete Evaluation Rule

Permanently deletes an evaluation rule. Metric results it already produced are kept; only the rule stops running. This action cannot be undone.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.evaluation_rules.delete(
    evaluation_rule_id="<EVALUATION-RULE-ID>",
)

For async mode, call a_delete and await it as shown below:

result = await client.evaluation_rules.a_delete(...)

Parameters

ParameterTypeDescription
evaluation_rule_idstrRequired. The id of the evaluation rule.

Returns

This method returns an object of type EvaluationRuleRef.

Types

EvaluationRule

A standing rule that runs a metric collection against matching production traces, spans or threads as they arrive.

class EvaluationRule:
    id: str
    name: str
    description: Optional[str]
    enabled: bool
    sample_rate: float = Field(alias="sampleRate")
    data_model: EvaluationRuleDataModel = Field(alias="dataModel")
    span_type: Optional[SpanType] = Field(alias="spanType")
    filters: Optional[FilterSet]
    thread_timelimit: Optional[int] = Field(alias="threadTimelimit")
    overwrite_evals: bool = Field(alias="overwriteEvals")
    metric_collection_id: str = Field(alias="metricCollectionId")
    created_at: str = Field(alias="createdAt")
    updated_at: str = Field(alias="updatedAt")

idstrRequired

The id of the rule, generated by Confident AI.

Example: "<EVALUATION-RULE-ID>"

namestrRequired

The name of the rule.

Example: "Score production answers"

descriptionOptional[str]Required

A note about what the rule checks, or null when unset.

Example: "Scores answers we serve to end users."

enabledboolRequired

Whether the rule is currently evaluating.

Example: true

sample_ratefloatRequired

The fraction of matching items the rule evaluates, between 0 and 1.

Example: 0.2

data_modelEvaluationRuleDataModelRequired

span_typeOptional[SpanType]Required

The kind of span the rule is narrowed to, or null when it evaluates every span or is not a SPAN rule.

See SpanType.

filtersOptional[FilterSet]Required

The filters an item must match to be evaluated, or null when the rule evaluates everything its dataModel covers.

See FilterSet.

Example: {"operator":"AND","groups":[{"operator":"AND","filters":[{"category":"Name","condition":"Is","value":"capital-lookup"}]}]}

thread_timelimitOptional[int]Required

For THREAD rules, the seconds of inactivity waited before a thread is evaluated, or null when the rule uses the project's thread timelimit.

Example: 300

overwrite_evalsboolRequired

Whether items that already have results for this metric collection are re-evaluated.

Example: false

metric_collection_idstrRequired

The id of the metric collection the rule runs. Retrieve it from the metric collections endpoint to see the metrics it holds.

Example: "<METRIC-COLLECTION-ID>"

created_atstrRequired

When the rule was created, as an ISO 8601 datetime.

Example: "2025-01-15T10:30:00+00:00"

updated_atstrRequired

When the rule was last changed, as an ISO 8601 datetime.

Example: "2025-01-20T08:00:00+00:00"

EvaluationRuleDataModel

What kind of production item a rule evaluates: TRACE for a whole trace, SPAN for a single step within one, THREAD for a finished conversation.

class EvaluationRuleDataModel(Enum):
    TRACE = "TRACE"
    SPAN = "SPAN"
    THREAD = "THREAD"

TRACE · SPAN · THREAD

EvaluationRuleList

One page of evaluation rules, with the total across all pages.

class EvaluationRuleList:
    evaluation_rules: List[EvaluationRuleSummary] = Field(alias="evaluationRules")
    total_evaluation_rules: int = Field(alias="totalEvaluationRules")
    page: int
    page_size: int = Field(alias="pageSize")

evaluation_rulesList[EvaluationRuleSummary]Required

The rules for the current page, newest first.

See EvaluationRuleSummary.

total_evaluation_rulesintRequired

The total number of rules matching the query, across all pages.

Example: 4

pageintRequired

The page this response covers.

Example: 1

page_sizeintRequired

The number of rules per page.

Example: 25

EvaluationRuleRef

A reference to an evaluation rule by its id.

class EvaluationRuleRef:
    id: str

idstrRequired

The id of the rule, generated by Confident AI.

Example: "<EVALUATION-RULE-ID>"

EvaluationRuleSummary

A rule as it appears in a list: enough to pick one. Its full configuration and the metric collection it runs come from retrieving it by id.

class EvaluationRuleSummary:
    id: str
    name: str
    enabled: bool
    data_model: EvaluationRuleDataModel = Field(alias="dataModel")

idstrRequired

The id of the rule, generated by Confident AI.

Example: "<EVALUATION-RULE-ID>"

namestrRequired

The name of the rule.

Example: "Score production answers"

enabledboolRequired

Whether the rule is currently evaluating.

Example: true

data_modelEvaluationRuleDataModelRequired

FilterSet

A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.

class FilterSet:
    operator: Literal["AND", "OR"]
    groups: List[FilterSetGroup]

operatorLiteral["AND", "OR"]Required

groupsList[FilterSetGroup]Required

SpanType

The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.

class SpanType(Enum):
    SPAN = "SPAN"
    AGENT = "AGENT"
    TOOL = "TOOL"
    RETRIEVER = "RETRIEVER"
    LLM = "LLM"

SPAN · AGENT · TOOL · RETRIEVER · LLM

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI