Evaluation Rules
Overview
The Confident AI SDK exposes every Evaluation Rule method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
List Evaluation Rules
Lists the evaluation rules in your Confident AI project one page at a time, newest first, as summary rows. Requires the Starter plan or above.
from confident_ai import ConfidentAI
from confident_ai.evaluation_rules import EvaluationRuleDataModel
client = ConfidentAI()
result = client.evaluation_rules.list(
data_model=EvaluationRuleDataModel.TRACE,
page=1,
page_size=25,
)For async mode, call a_list and await it as shown below:
result = await client.evaluation_rules.a_list(...)Parameters
| Parameter | Type | Description |
|---|---|---|
data_model | Optional[EvaluationRuleDataModel] | See EvaluationRuleDataModel. |
page | Optional[int] | The page to return. Defaults to 1. |
page_size | Optional[int] | The number of results per page, at most 100. Defaults to 25. |
import { ConfidentAI } from "confident-ai";
import { EvaluationRuleDataModel } from "confident-ai/evaluation-rules";
const client = new ConfidentAI();
const result = await client.evaluationRules.list(
{ dataModel: EvaluationRuleDataModel.TRACE, page: 1, pageSize: 25 },
);Parameters
| Parameter | Type | Description |
|---|---|---|
dataModel | EvaluationRuleDataModel | See EvaluationRuleDataModel. |
page | number | The page to return. Defaults to 1. |
pageSize | number | The number of results per page, at most 100. Defaults to 25. |
Returns
This method returns an object of type EvaluationRuleList.
Create Evaluation Rule
Creates a standing rule that runs a metric collection against matching production traces, spans or threads as they arrive, and returns its id. Running metrics consumes LLM usage. The collection's turn type must match the rule — THREAD rules need a multi-turn collection, TRACE and SPAN rules a single-turn one — and only one enabled THREAD rule may target a given collection. Requires the Starter plan or above.
from confident_ai import ConfidentAI
from confident_ai.evaluation_rules import EvaluationRuleDataModel
from confident_ai.common import SpanType
client = ConfidentAI()
result = client.evaluation_rules.create(
name="Score production answers",
data_model=EvaluationRuleDataModel.TRACE,
metric_collection_id="<METRIC-COLLECTION-ID>",
enabled=True,
description="Scores answers we serve to end users.",
sample_rate=0.2,
span_type=SpanType.SPAN,
filters={
"operator": "AND",
"groups": [
{
"operator": "AND",
"filters": [
{
"category": "Name",
"condition": "Is",
"value": "capital-lookup"
}
]
}
]
},
thread_timelimit=300,
overwrite_evals=False,
)For async mode, call a_create and await it as shown below:
result = await client.evaluation_rules.a_create(...)Parameters
| Parameter | Type | Description |
|---|---|---|
name | str | Required. A name for the rule, unique within the project. |
data_model | EvaluationRuleDataModel | Required. See EvaluationRuleDataModel. |
metric_collection_id | str | Required. The id of the metric collection to run. It must be multi-turn for THREAD rules and single-turn for TRACE and SPAN rules. |
enabled | Optional[bool] | Whether the rule evaluates matching items as they arrive. Defaults to true. |
description | Optional[str] | A note about what the rule checks. Send null to clear it. |
sample_rate | Optional[float] | The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them. |
span_type | Optional[SpanType] | Only evaluate spans of this kind. Allowed only when dataModel is SPAN, and cleared automatically if the rule moves off SPAN. Send null to evaluate every span. See SpanType. |
filters | Optional[FilterSet] | Only evaluate items matching these filters. Send null to evaluate every item the rule's dataModel covers. See FilterSet. |
thread_timelimit | Optional[int] | For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300. |
overwrite_evals | Optional[bool] | Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false. |
import { ConfidentAI } from "confident-ai";
import { SpanType } from "confident-ai/common";
import { EvaluationRuleDataModel } from "confident-ai/evaluation-rules";
const client = new ConfidentAI();
const result = await client.evaluationRules.create(
"Score production answers",
EvaluationRuleDataModel.TRACE,
"<METRIC-COLLECTION-ID>",
{
enabled: true,
description: "Scores answers we serve to end users.",
sampleRate: 0.2,
spanType: SpanType.SPAN,
filters: {
operator: "AND",
groups: [
{
operator: "AND",
filters: [{ category: "Name", condition: "Is", value: "capital-lookup" }]
}
]
},
threadTimelimit: 300,
overwriteEvals: false
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
name | string | Required. A name for the rule, unique within the project. |
dataModel | EvaluationRuleDataModel | Required. See EvaluationRuleDataModel. |
metricCollectionId | string | Required. The id of the metric collection to run. It must be multi-turn for THREAD rules and single-turn for TRACE and SPAN rules. |
enabled | boolean | Whether the rule evaluates matching items as they arrive. Defaults to true. |
description | string | null | A note about what the rule checks. Send null to clear it. |
sampleRate | number | The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them. |
spanType | SpanType | null | Only evaluate spans of this kind. Allowed only when dataModel is SPAN, and cleared automatically if the rule moves off SPAN. Send null to evaluate every span. See SpanType. |
filters | FilterSet | null | Only evaluate items matching these filters. Send null to evaluate every item the rule's dataModel covers. See FilterSet. |
threadTimelimit | number | null | For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300. |
overwriteEvals | boolean | Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false. |
Returns
This method returns an object of type EvaluationRuleRef.
Get Evaluation Rule
Retrieves an evaluation rule by id with its full configuration, including the filters an item must match and the id of the metric collection it runs.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.evaluation_rules.get(
evaluation_rule_id="<EVALUATION-RULE-ID>",
)For async mode, call a_get and await it as shown below:
result = await client.evaluation_rules.a_get(...)Parameters
| Parameter | Type | Description |
|---|---|---|
evaluation_rule_id | str | Required. The id of the evaluation rule. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.evaluationRules.get("<EVALUATION-RULE-ID>");Parameters
| Parameter | Type | Description |
|---|---|---|
evaluationRuleId | string | Required. The id of the evaluation rule. |
Returns
This method returns an object of type EvaluationRule.
Update Evaluation Rule
Updates an evaluation rule and returns it. Only the fields you send are changed; omitting a field leaves it untouched, and sending null clears it. Constraints are re-checked against the rule the update produces, not just the fields you sent, so switching a rule to THREAD still requires a multi-turn metric collection.
from confident_ai import ConfidentAI
from confident_ai.evaluation_rules import EvaluationRuleDataModel
from confident_ai.common import SpanType
client = ConfidentAI()
result = client.evaluation_rules.update(
evaluation_rule_id="<EVALUATION-RULE-ID>",
name="Score production answers",
enabled=False,
data_model=EvaluationRuleDataModel.TRACE,
metric_collection_id="<METRIC-COLLECTION-ID>",
description="Scores answers we serve to end users.",
sample_rate=0.2,
span_type=SpanType.SPAN,
filters={
"operator": "AND",
"groups": [
{
"operator": "AND",
"filters": [
{
"category": "Name",
"condition": "Is",
"value": "capital-lookup"
}
]
}
]
},
thread_timelimit=300,
overwrite_evals=False,
)For async mode, call a_update and await it as shown below:
result = await client.evaluation_rules.a_update(...)Parameters
| Parameter | Type | Description |
|---|---|---|
evaluation_rule_id | str | Required. The id of the evaluation rule. |
name | Optional[str] | A new name for the rule, unique within the project. |
enabled | Optional[bool] | Whether the rule evaluates matching items as they arrive. |
data_model | Optional[EvaluationRuleDataModel] | See EvaluationRuleDataModel. |
metric_collection_id | Optional[str] | The id of a different metric collection to run. |
description | Optional[str] | A note about what the rule checks. Send null to clear it. |
sample_rate | Optional[float] | The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them. |
span_type | Optional[SpanType] | Only evaluate spans of this kind. Allowed only when dataModel is SPAN, and cleared automatically if the rule moves off SPAN. Send null to evaluate every span. See SpanType. |
filters | Optional[FilterSet] | Only evaluate items matching these filters. Send null to evaluate every item the rule's dataModel covers. See FilterSet. |
thread_timelimit | Optional[int] | For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300. |
overwrite_evals | Optional[bool] | Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false. |
import { ConfidentAI } from "confident-ai";
import { SpanType } from "confident-ai/common";
import { EvaluationRuleDataModel } from "confident-ai/evaluation-rules";
const client = new ConfidentAI();
const result = await client.evaluationRules.update(
"<EVALUATION-RULE-ID>",
{
name: "Score production answers",
enabled: false,
dataModel: EvaluationRuleDataModel.TRACE,
metricCollectionId: "<METRIC-COLLECTION-ID>",
description: "Scores answers we serve to end users.",
sampleRate: 0.2,
spanType: SpanType.SPAN,
filters: {
operator: "AND",
groups: [
{
operator: "AND",
filters: [{ category: "Name", condition: "Is", value: "capital-lookup" }]
}
]
},
threadTimelimit: 300,
overwriteEvals: false
},
);Parameters
| Parameter | Type | Description |
|---|---|---|
evaluationRuleId | string | Required. The id of the evaluation rule. |
name | string | A new name for the rule, unique within the project. |
enabled | boolean | Whether the rule evaluates matching items as they arrive. |
dataModel | EvaluationRuleDataModel | See EvaluationRuleDataModel. |
metricCollectionId | string | The id of a different metric collection to run. |
description | string | null | A note about what the rule checks. Send null to clear it. |
sampleRate | number | The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them. |
spanType | SpanType | null | Only evaluate spans of this kind. Allowed only when dataModel is SPAN, and cleared automatically if the rule moves off SPAN. Send null to evaluate every span. See SpanType. |
filters | FilterSet | null | Only evaluate items matching these filters. Send null to evaluate every item the rule's dataModel covers. See FilterSet. |
threadTimelimit | number | null | For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300. |
overwriteEvals | boolean | Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false. |
Returns
This method returns an object of type EvaluationRule.
Delete Evaluation Rule
Permanently deletes an evaluation rule. Metric results it already produced are kept; only the rule stops running. This action cannot be undone.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.evaluation_rules.delete(
evaluation_rule_id="<EVALUATION-RULE-ID>",
)For async mode, call a_delete and await it as shown below:
result = await client.evaluation_rules.a_delete(...)Parameters
| Parameter | Type | Description |
|---|---|---|
evaluation_rule_id | str | Required. The id of the evaluation rule. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.evaluationRules.delete("<EVALUATION-RULE-ID>");Parameters
| Parameter | Type | Description |
|---|---|---|
evaluationRuleId | string | Required. The id of the evaluation rule. |
Returns
This method returns an object of type EvaluationRuleRef.
Types
EvaluationRule
A standing rule that runs a metric collection against matching production traces, spans or threads as they arrive.
class EvaluationRule:
id: str
name: str
description: Optional[str]
enabled: bool
sample_rate: float = Field(alias="sampleRate")
data_model: EvaluationRuleDataModel = Field(alias="dataModel")
span_type: Optional[SpanType] = Field(alias="spanType")
filters: Optional[FilterSet]
thread_timelimit: Optional[int] = Field(alias="threadTimelimit")
overwrite_evals: bool = Field(alias="overwriteEvals")
metric_collection_id: str = Field(alias="metricCollectionId")
created_at: str = Field(alias="createdAt")
updated_at: str = Field(alias="updatedAt")idstrRequired
The id of the rule, generated by Confident AI.
Example: "<EVALUATION-RULE-ID>"
namestrRequired
The name of the rule.
Example: "Score production answers"
descriptionOptional[str]Required
A note about what the rule checks, or null when unset.
Example: "Scores answers we serve to end users."
enabledboolRequired
Whether the rule is currently evaluating.
Example: true
sample_ratefloatRequired
The fraction of matching items the rule evaluates, between 0 and 1.
Example: 0.2
data_modelEvaluationRuleDataModelRequired
span_typeOptional[SpanType]Required
The kind of span the rule is narrowed to, or null when it evaluates every span or is not a SPAN rule.
See SpanType.
filtersOptional[FilterSet]Required
The filters an item must match to be evaluated, or null when the rule evaluates everything its dataModel covers.
See FilterSet.
Example: {"operator":"AND","groups":[{"operator":"AND","filters":[{"category":"Name","condition":"Is","value":"capital-lookup"}]}]}
thread_timelimitOptional[int]Required
For THREAD rules, the seconds of inactivity waited before a thread is evaluated, or null when the rule uses the project's thread timelimit.
Example: 300
overwrite_evalsboolRequired
Whether items that already have results for this metric collection are re-evaluated.
Example: false
metric_collection_idstrRequired
The id of the metric collection the rule runs. Retrieve it from the metric collections endpoint to see the metrics it holds.
Example: "<METRIC-COLLECTION-ID>"
created_atstrRequired
When the rule was created, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
updated_atstrRequired
When the rule was last changed, as an ISO 8601 datetime.
Example: "2025-01-20T08:00:00+00:00"
interface EvaluationRule {
id: string;
name: string;
description: string | null;
enabled: boolean;
sampleRate: number;
dataModel: EvaluationRuleDataModel;
spanType: SpanType | null;
filters: FilterSet | null;
threadTimelimit: number | null;
overwriteEvals: boolean;
metricCollectionId: string;
createdAt: string;
updatedAt: string;
}idstringRequired
The id of the rule, generated by Confident AI.
Example: "<EVALUATION-RULE-ID>"
namestringRequired
The name of the rule.
Example: "Score production answers"
descriptionstring | nullRequired
A note about what the rule checks, or null when unset.
Example: "Scores answers we serve to end users."
enabledbooleanRequired
Whether the rule is currently evaluating.
Example: true
sampleRatenumberRequired
The fraction of matching items the rule evaluates, between 0 and 1.
Example: 0.2
dataModelEvaluationRuleDataModelRequired
spanTypeSpanType | nullRequired
The kind of span the rule is narrowed to, or null when it evaluates every span or is not a SPAN rule.
See SpanType.
filtersFilterSet | nullRequired
The filters an item must match to be evaluated, or null when the rule evaluates everything its dataModel covers.
See FilterSet.
Example: {"operator":"AND","groups":[{"operator":"AND","filters":[{"category":"Name","condition":"Is","value":"capital-lookup"}]}]}
threadTimelimitnumber | nullRequired
For THREAD rules, the seconds of inactivity waited before a thread is evaluated, or null when the rule uses the project's thread timelimit.
Example: 300
overwriteEvalsbooleanRequired
Whether items that already have results for this metric collection are re-evaluated.
Example: false
metricCollectionIdstringRequired
The id of the metric collection the rule runs. Retrieve it from the metric collections endpoint to see the metrics it holds.
Example: "<METRIC-COLLECTION-ID>"
createdAtstringRequired
When the rule was created, as an ISO 8601 datetime.
Example: "2025-01-15T10:30:00+00:00"
updatedAtstringRequired
When the rule was last changed, as an ISO 8601 datetime.
Example: "2025-01-20T08:00:00+00:00"
EvaluationRuleDataModel
What kind of production item a rule evaluates: TRACE for a whole trace, SPAN for a single step within one, THREAD for a finished conversation.
class EvaluationRuleDataModel(Enum):
TRACE = "TRACE"
SPAN = "SPAN"
THREAD = "THREAD"enum EvaluationRuleDataModel {
TRACE = "TRACE",
SPAN = "SPAN",
THREAD = "THREAD",
}TRACE · SPAN · THREAD
EvaluationRuleList
One page of evaluation rules, with the total across all pages.
class EvaluationRuleList:
evaluation_rules: List[EvaluationRuleSummary] = Field(alias="evaluationRules")
total_evaluation_rules: int = Field(alias="totalEvaluationRules")
page: int
page_size: int = Field(alias="pageSize")evaluation_rulesList[EvaluationRuleSummary]Required
The rules for the current page, newest first.
total_evaluation_rulesintRequired
The total number of rules matching the query, across all pages.
Example: 4
pageintRequired
The page this response covers.
Example: 1
page_sizeintRequired
The number of rules per page.
Example: 25
interface EvaluationRuleList {
evaluationRules: EvaluationRuleSummary[];
totalEvaluationRules: number;
page: number;
pageSize: number;
}evaluationRulesEvaluationRuleSummary[]Required
The rules for the current page, newest first.
totalEvaluationRulesnumberRequired
The total number of rules matching the query, across all pages.
Example: 4
pagenumberRequired
The page this response covers.
Example: 1
pageSizenumberRequired
The number of rules per page.
Example: 25
EvaluationRuleRef
A reference to an evaluation rule by its id.
class EvaluationRuleRef:
id: stridstrRequired
The id of the rule, generated by Confident AI.
Example: "<EVALUATION-RULE-ID>"
interface EvaluationRuleRef {
id: string;
}idstringRequired
The id of the rule, generated by Confident AI.
Example: "<EVALUATION-RULE-ID>"
EvaluationRuleSummary
A rule as it appears in a list: enough to pick one. Its full configuration and the metric collection it runs come from retrieving it by id.
class EvaluationRuleSummary:
id: str
name: str
enabled: bool
data_model: EvaluationRuleDataModel = Field(alias="dataModel")idstrRequired
The id of the rule, generated by Confident AI.
Example: "<EVALUATION-RULE-ID>"
namestrRequired
The name of the rule.
Example: "Score production answers"
enabledboolRequired
Whether the rule is currently evaluating.
Example: true
data_modelEvaluationRuleDataModelRequired
interface EvaluationRuleSummary {
id: string;
name: string;
enabled: boolean;
dataModel: EvaluationRuleDataModel;
}idstringRequired
The id of the rule, generated by Confident AI.
Example: "<EVALUATION-RULE-ID>"
namestringRequired
The name of the rule.
Example: "Score production answers"
enabledbooleanRequired
Whether the rule is currently evaluating.
Example: true
dataModelEvaluationRuleDataModelRequired
FilterSet
A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.
class FilterSet:
operator: Literal["AND", "OR"]
groups: List[FilterSetGroup]operatorLiteral["AND", "OR"]Required
groupsList[FilterSetGroup]Required
interface FilterSet {
operator: "AND" | "OR";
groups: FilterSetGroup[];
}operator"AND" | "OR"Required
groupsFilterSetGroup[]Required
SpanType
The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.
class SpanType(Enum):
SPAN = "SPAN"
AGENT = "AGENT"
TOOL = "TOOL"
RETRIEVER = "RETRIEVER"
LLM = "LLM"enum SpanType {
SPAN = "SPAN",
AGENT = "AGENT",
TOOL = "TOOL",
RETRIEVER = "RETRIEVER",
LLM = "LLM",
}SPAN · AGENT · TOOL · RETRIEVER · LLM
Last updated on