Launch Week 3: Five days of launches

Versions

Overview

The Confident AI SDK exposes every Version method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Governance Control Versions

Lists a governance control's definition history, newest first, so the first entry is the rule it evaluates today. Versions are append-only, which makes this the record of how the check has changed — pass a version's version label to the assessments endpoint to read the verdicts it produced.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.organization.list_governance_control_versions(
    control_id="<GOVERNANCE-CONTROL-ID>",
    page=1,
    page_size=25,
)

For async mode, call a_list_governance_control_versions and await it as shown below:

result = await client.organization.a_list_governance_control_versions(...)

Parameters

ParameterTypeDescription
control_idstrRequired. The id of the governance control.
pageOptional[int]The page to return. Defaults to 1.
page_sizeOptional[int]The number of versions per page, at most 100. Defaults to 25.

Returns

This method returns an object of type GovernanceControlVersionList.

Governance Control Runtime Config

Changes what a governance control checks by appending a new version of its definition. The version that was current stays in the history with the verdicts computed against it, and the new one becomes what the next assessment runs against; existing verdicts are never recomputed or migrated. Which request shape is expected follows the control's own type — a pre-deployment control takes the pre-deployment config, every other type the runtime config — and every field is required even when null, so send the definition in full rather than a patch.

The rule a RUNTIME control evaluates, read as one sentence: aggregate aggregation over dataModel for the trailing 24 hours, restricted to filters, and fail when the result sits on the direction side of the threshold. Aggregating Error rate over TRACE against a threshold of 0.02 above fails a project whose traces errored on more than 2% of requests in the last day. The window is fixed and is not part of the definition. The aggregation has to be one the data model supports.

from confident_ai import ConfidentAI
from confident_ai.common import FilterSet
from confident_ai.organization import GovernanceControlAggregation
from confident_ai.organization import GovernanceControlDataModel
from confident_ai.organization import GovernanceControlExtraQueryParams
from confident_ai.organization import GovernanceControlMetricDataCategory
from confident_ai.organization import GovernanceControlRuntimeConfig
from confident_ai.organization import GovernanceControlSeverity
from confident_ai.organization import GovernanceControlThresholdDirection
from confident_ai.organization import GovernanceControlThresholdSettings

client = ConfidentAI()

result = client.organization.create_governance_control_version(
    control_id="<GOVERNANCE-CONTROL-ID>",
    control_config=GovernanceControlRuntimeConfig(
        data_model=GovernanceControlDataModel.TRACE,
        aggregation=GovernanceControlAggregation.AVG_COST,
        threshold_settings=GovernanceControlThresholdSettings(
            value=0.02,
            direction=GovernanceControlThresholdDirection.ABOVE
        ),
        extra_query_params=GovernanceControlExtraQueryParams(
            category=GovernanceControlMetricDataCategory.TRACE
        ),
        filters=FilterSet(
            operator="AND",
            groups=[]
        ),
        severity=GovernanceControlSeverity.CRITICAL
    ),
)

For async mode, call a_create_governance_control_version and await it as shown below:

result = await client.organization.a_create_governance_control_version(...)

Parameters

ParameterTypeDescription
control_idstrRequired. The id of the governance control.
control_configCreateGovernanceControlVersionRequestRequired. The definition to snapshot as the control's next version. Which of the two shapes is expected is decided by the control's own type, not by what you send: a pre-deployment control takes the pre-deployment config, and every other type — RUNTIME and OPERATIONAL alike — takes the runtime config. Except for extraQueryParams, every field is required even when null, so a version is a complete definition rather than a patch of the one before it. Pass a GovernanceControlRuntimeConfig or a GovernanceControlPreDeploymentConfig. See CreateGovernanceControlVersionRequest.

Returns

This method returns an object of type GovernanceControlVersion.

Governance Control Pre Deployment Config

Changes what a governance control checks by appending a new version of its definition. The version that was current stays in the history with the verdicts computed against it, and the new one becomes what the next assessment runs against; existing verdicts are never recomputed or migrated. Which request shape is expected follows the control's own type — a pre-deployment control takes the pre-deployment config, every other type the runtime config — and every field is required even when null, so send the definition in full rather than a patch.

The rule a PRE_DEPLOYMENT_EVALS or PRE_DEPLOYMENT_RED_TEAMING control evaluates, read as one sentence: find the project's newest completed run whose identifier is identifier within the last window.days days, and pass when that run satisfies filters. An identifier of pre-release with a 30-day window and a filter of Pass rate >= 0.9 fails a project whose last pre-release run scored below 90%, and reports NO_DATA when it has not run one at all. PRE_DEPLOYMENT_EVALS looks at test runs and PRE_DEPLOYMENT_RED_TEAMING at red teaming runs; the endpoint picks which by the control's own type, so the same shape serves both. Send a non-empty identifier unless officialOnly is true.

from confident_ai import ConfidentAI
from confident_ai.common import FilterSet
from confident_ai.organization import GovernanceControlPreDeploymentConfig
from confident_ai.organization import GovernanceControlPreDeploymentWindow
from confident_ai.organization import GovernanceControlSeverity

client = ConfidentAI()

result = client.organization.create_governance_control_version(
    control_id="<GOVERNANCE-CONTROL-ID>",
    control_config=GovernanceControlPreDeploymentConfig(
        identifier="pre-release",
        window=GovernanceControlPreDeploymentWindow(
            days=30
        ),
        official_only=False,
        filters=FilterSet(
            operator="AND",
            groups=[]
        ),
        severity=GovernanceControlSeverity.CRITICAL
    ),
)

For async mode, call a_create_governance_control_version and await it as shown below:

result = await client.organization.a_create_governance_control_version(...)

Parameters

ParameterTypeDescription
control_idstrRequired. The id of the governance control.
control_configCreateGovernanceControlVersionRequestRequired. The definition to snapshot as the control's next version. Which of the two shapes is expected is decided by the control's own type, not by what you send: a pre-deployment control takes the pre-deployment config, and every other type — RUNTIME and OPERATIONAL alike — takes the runtime config. Except for extraQueryParams, every field is required even when null, so a version is a complete definition rather than a patch of the one before it. Pass a GovernanceControlRuntimeConfig or a GovernanceControlPreDeploymentConfig. See CreateGovernanceControlVersionRequest.

Returns

This method returns an object of type GovernanceControlVersion.

Types

CreateGovernanceControlVersionRequest

The definition to snapshot as the control's next version. Which of the two shapes is expected is decided by the control's own type, not by what you send: a pre-deployment control takes the pre-deployment config, and every other type — RUNTIME and OPERATIONAL alike — takes the runtime config. Except for extraQueryParams, every field is required even when null, so a version is a complete definition rather than a patch of the one before it.

CreateGovernanceControlVersionRequest = Union[
    GovernanceControlRuntimeConfig,
    GovernanceControlPreDeploymentConfig,
]

A CreateGovernanceControlVersionRequest is one of the shapes below. Send the fields of one of them, never a mix of both.

The rule a RUNTIME control evaluates, read as one sentence: aggregate aggregation over dataModel for the trailing 24 hours, restricted to filters, and fail when the result sits on the direction side of the threshold. Aggregating Error rate over TRACE against a threshold of 0.02 above fails a project whose traces errored on more than 2% of requests in the last day. The window is fixed and is not part of the definition. The aggregation has to be one the data model supports.

class GovernanceControlRuntimeConfig:
    data_model: Optional[GovernanceControlDataModel] = Field(alias="dataModel")
    aggregation: Optional[GovernanceControlAggregation]
    threshold_settings: Optional[GovernanceControlThresholdSettings] = Field(alias="thresholdSettings")
    extra_query_params: Optional[GovernanceControlExtraQueryParams] = Field(default=None, alias="extraQueryParams")
    filters: Optional[FilterSet]
    severity: Optional[GovernanceControlSeverity]

data_modelOptional[GovernanceControlDataModel]Required

The production data to measure. Send null to leave the control unconfigured, which makes it assess as ERROR until it is set.

See GovernanceControlDataModel.

aggregationOptional[GovernanceControlAggregation]Required

How to reduce the measured data to the one number the threshold is compared against. It must be an aggregation the selected dataModel supports.

See GovernanceControlAggregation.

threshold_settingsOptional[GovernanceControlThresholdSettings]Required

The threshold the aggregated value is compared against. Send null to leave the control unconfigured.

See GovernanceControlThresholdSettings.

extra_query_paramsOptional[GovernanceControlExtraQueryParams]

Extra scoping for the measured data. It is the one field that is carried over from the current version when you omit it; send null to clear it.

See GovernanceControlExtraQueryParams.

filtersOptional[FilterSet]Required

Narrows the data that is aggregated, so the control measures a slice of production rather than all of it. Send null to measure everything.

See FilterSet.

severityOptional[GovernanceControlSeverity]Required

How much a failure matters. Send null to leave it unset, which still blocks a deployment gate.

See GovernanceControlSeverity.

FilterSet

A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.

class FilterSet:
    operator: Literal["AND", "OR"]
    groups: List[FilterSetGroup]

operatorLiteral["AND", "OR"]Required

groupsList[FilterSetGroup]Required

GovernanceControlAggregation

How a runtime control reduces the data it measures to the single number it compares against its threshold. Each aggregation is only valid for some data models — Avg score and Pass rate need METRIC_DATA, Avg rating needs ANNOTATION, Input tokens needs SPAN — and a pairing the selected dataModel does not support is rejected.

class GovernanceControlAggregation(Enum):
    AVG_COST = "Avg cost"
    AVG_LATENCY = "Avg latency"
    AVG_RATING = "Avg rating"
    AVG_SCORE = "Avg score"
    AVG_VALUE = "Avg value"
    COUNT = "Count"
    ERROR_COUNT = "Error count"
    ERROR_RATE = "Error rate"
    FAILURE_RATE = "Failure rate"
    INPUT_COST = "Input cost"
    INPUT_TOKENS = "Input tokens"
    MEDIAN_SCORE = "Median score"
    OUTPUT_COST = "Output cost"
    OUTPUT_TOKENS = "Output tokens"
    P50_LATENCY = "P50 latency"
    P90_LATENCY = "P90 latency"
    P99_LATENCY = "P99 latency"
    PASS_RATE = "Pass rate"
    TOTAL_COST = "Total cost"
    TOTAL_TOKENS = "Total tokens"
    UNIQUE_END_USERS = "Unique end users"
    UNIQUE_METADATA_VALUES = "Unique metadata values"
    UNIQUE_THREADS = "Unique threads"

AVG_COST · AVG_LATENCY · AVG_RATING · AVG_SCORE · AVG_VALUE · COUNT · ERROR_COUNT · ERROR_RATE · FAILURE_RATE · INPUT_COST · INPUT_TOKENS · MEDIAN_SCORE · OUTPUT_COST · OUTPUT_TOKENS · P50_LATENCY · P90_LATENCY · P99_LATENCY · PASS_RATE · TOTAL_COST · TOTAL_TOKENS · UNIQUE_END_USERS · UNIQUE_METADATA_VALUES · UNIQUE_THREADS

GovernanceControlDataModel

The production data a runtime control measures: TRACE for whole requests, SPAN for individual steps, THREAD for conversations, METRIC_DATA for evaluation scores, and ANNOTATION for human ratings. It decides which aggregations are valid.

class GovernanceControlDataModel(Enum):
    TRACE = "TRACE"
    SPAN = "SPAN"
    THREAD = "THREAD"
    METRIC_DATA = "METRIC_DATA"
    ANNOTATION = "ANNOTATION"

TRACE · SPAN · THREAD · METRIC_DATA · ANNOTATION

GovernanceControlExtraQueryParams

Extra scoping for the data a runtime control measures, beyond its data model and filters.

class GovernanceControlExtraQueryParams:
    category: Optional[GovernanceControlMetricDataCategory] = None

categoryOptional[GovernanceControlMetricDataCategory]

GovernanceControlMetricDataCategory

Which kind of item the evaluation scores were recorded on, for a control measuring METRIC_DATA.

class GovernanceControlMetricDataCategory(Enum):
    TRACE = "TRACE"
    SPAN = "SPAN"
    THREAD = "THREAD"

TRACE · SPAN · THREAD

GovernanceControlPreDeploymentSettings

Which run a pre-deployment control gates on. The fields are individually optional because a version snapshotted before the control was configured stores an empty object; a configured control always carries either identifier and window or officialOnly set to true.

class GovernanceControlPreDeploymentSettings:
    identifier: Optional[str] = None
    window: Optional[GovernanceControlPreDeploymentWindow] = None
    official_only: Optional[bool] = Field(default=None, alias="officialOnly")

identifierOptional[str]

The identifier of the test run or red teaming run the control gates on. Absent on a control that gates on the project's official run instead.

Example: "pre-release"

windowOptional[GovernanceControlPreDeploymentWindow]

official_onlyOptional[bool]

Whether the control gates on the project's most recent official run rather than on a run matching identifier.

Example: false

GovernanceControlPreDeploymentWindow

The rolling lookback a pre-deployment control searches for the run it gates on.

class GovernanceControlPreDeploymentWindow:
    days: int

daysintRequired

How many days back the control looks for a run. It is stored as a day count rather than as dates, so the gate does not go stale as it is re-assessed.

Example: 30

GovernanceControlSeverity

How much a failing control matters, set per version rather than per control. LOW never blocks a deployment gate; CRITICAL, HIGH and MEDIUM block, and so does leaving the severity unset.

class GovernanceControlSeverity(Enum):
    CRITICAL = "CRITICAL"
    HIGH = "HIGH"
    MEDIUM = "MEDIUM"
    LOW = "LOW"

CRITICAL · HIGH · MEDIUM · LOW

GovernanceControlStoredThresholdSettings

A threshold as a version stores it. Both fields are set on a configured runtime control; either can be absent on a version snapshotted before the control was configured, which is what makes it unconfigured.

class GovernanceControlStoredThresholdSettings:
    value: Optional[float] = None
    direction: Optional[GovernanceControlThresholdDirection] = None

valueOptional[float]

The number the aggregated value is compared against, in the unit the aggregation produces — a rate is a fraction between 0 and 1, a latency is in milliseconds, a cost is in USD.

Example: 0.02

directionOptional[GovernanceControlThresholdDirection]

GovernanceControlThresholdDirection

Which side of the threshold fails: above fails once the measured value rises past value, below fails once it drops under it.

class GovernanceControlThresholdDirection(Enum):
    ABOVE = "above"
    BELOW = "below"

ABOVE · BELOW

GovernanceControlThresholdSettings

The comparison that turns a runtime control's measured value into a verdict.

class GovernanceControlThresholdSettings:
    value: float
    direction: GovernanceControlThresholdDirection

valuefloatRequired

The number the aggregated value is compared against, in the unit the aggregation produces — a rate is a fraction between 0 and 1, a latency is in milliseconds, a cost is in USD.

Example: 0.02

directionGovernanceControlThresholdDirectionRequired

GovernanceControlVersion

An immutable snapshot of what a control checks. Which fields are populated follows the control's type: a runtime control carries dataModel, aggregation and thresholdSettings, a pre-deployment control carries preDeploymentSettings, and an OPERATIONAL control carries neither because its check ships with the platform. Editing a control appends a new version rather than changing this one, so every past verdict keeps pointing at the definition it was computed against.

class GovernanceControlVersion:
    id: str
    version: str
    sequence: int
    data_model: Optional[GovernanceControlDataModel] = Field(alias="dataModel")
    aggregation: Optional[str]
    threshold_settings: Optional[GovernanceControlStoredThresholdSettings] = Field(alias="thresholdSettings")
    extra_query_params: Optional[GovernanceControlExtraQueryParams] = Field(alias="extraQueryParams")
    pre_deployment_settings: Optional[GovernanceControlPreDeploymentSettings] = Field(alias="preDeploymentSettings")
    filters: Optional[FilterSet]
    severity: Optional[GovernanceControlSeverity]
    assessments_count: int = Field(alias="assessmentsCount")
    created_at: str = Field(alias="createdAt")

idstrRequired

The id of the version, generated by Confident AI.

Example: "<GOVERNANCE-CONTROL-VERSION-ID>"

versionstrRequired

The human-readable label for sequence, written as three two-digit groups that roll over at 100, so sequence 2 is 00.00.02 and sequence 100 is 00.01.00.

Example: "00.00.02"

sequenceintRequired

The position of this version in the control's history, counting from 1. The highest sequence is the current version, which is the one every new assessment runs against.

Example: 2

data_modelOptional[GovernanceControlDataModel]Required

The production data this version measures. It is set on a runtime control and null on every other type.

See GovernanceControlDataModel.

aggregationOptional[str]Required

How the measured data is reduced to one number, as one of the GovernanceControlAggregation values. It is set on a runtime control and null on every other type.

Example: "Error rate"

threshold_settingsOptional[GovernanceControlStoredThresholdSettings]Required

The threshold the aggregated value is compared against. It is set on a runtime control and null on every other type.

See GovernanceControlStoredThresholdSettings.

extra_query_paramsOptional[GovernanceControlExtraQueryParams]Required

Extra scoping for the measured data, or null when none was set. It is the one field carried over from the previous version when a new version omits it.

See GovernanceControlExtraQueryParams.

pre_deployment_settingsOptional[GovernanceControlPreDeploymentSettings]Required

Which run this version gates on. It is set on a pre-deployment control and null on every other type.

See GovernanceControlPreDeploymentSettings.

filtersOptional[FilterSet]Required

Narrows what the control looks at, or null when it looks at everything. On a runtime control the filters narrow the data that is aggregated; on a pre-deployment control they are matched against the run itself, so the run has to satisfy them for the control to pass. Versions written before filters took a { operator, groups } wrapper may still return a bare array of groups.

See FilterSet.

severityOptional[GovernanceControlSeverity]Required

How much a failure of this version matters, or null when it was left unset.

See GovernanceControlSeverity.

assessments_countintRequired

How many verdicts were recorded against this version.

Example: 64

created_atstrRequired

When this version was snapshotted.

Example: "2025-01-18T16:45:00+00:00"

GovernanceControlVersionList

One page of a control's definition history, with the total across all pages.

class GovernanceControlVersionList:
    versions: List[GovernanceControlVersion]
    total_governance_control_versions: int = Field(alias="totalGovernanceControlVersions")
    page: int
    page_size: int = Field(alias="pageSize")

versionsList[GovernanceControlVersion]Required

The control's definition history for the current page, newest version first, so the first entry of the first page is the current definition.

See GovernanceControlVersion.

total_governance_control_versionsintRequired

The number of versions this control has, across every page.

Example: 2

pageintRequired

The page this response covers.

Example: 1

page_sizeintRequired

The number of versions per page.

Example: 25

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI