Launch Week 3: Five days of launches

RT Frameworks

Every RT Frameworks method in the Confident AI Python and TypeScript SDKs.

Overview

The Confident AI SDK exposes every RT Framework method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List RT Frameworks

Lists the red teaming frameworks in your Confident AI project one page at a time, ordered by name. Each framework is returned with its risk categories counted; retrieve one by id to see what those categories select.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.rt_frameworks.list(page=1, page_size=25)

For async mode, call a_list and await it as shown below:

result = await client.rt_frameworks.a_list(...)

Parameters

ParameterTypeDescription
pageOptional[int]The page to return. Defaults to 1.
page_sizeOptional[int]The number of results per page, at most 100. Defaults to 25.

Returns

This method returns an object of type RTFrameworkList.

Create RT Framework

Creates a red teaming framework in your Confident AI project and returns its id. Send template to fill it from a Confident AI template rather than starting empty.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.rt_frameworks.create(
    name="OWASP Top 10 for LLMs",
    description="Our baseline coverage before each release.",
    template="EU Artificial Intelligence Act",
)

For async mode, call a_create and await it as shown below:

result = await client.rt_frameworks.a_create(...)

Parameters

ParameterTypeDescription
namestrRequired. The name of the framework, unique within the project.
descriptionOptional[str]What the framework covers. Send null to leave it unset.
templateOptional[str]A Confident AI template to fill the framework from, which creates its risk categories with vulnerability types and attack methods already selected. Omit it for an empty framework.

Returns

This method returns an object of type RTFrameworkRef.

Get RT Framework

Retrieves a red teaming framework by id, with each risk category resolved to the vulnerabilities it probes for and the attack methods it probes with, exactly as a run would use them.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.rt_frameworks.get(rt_framework_id="<RT-FRAMEWORK-ID>")

For async mode, call a_get and await it as shown below:

result = await client.rt_frameworks.a_get(...)

Parameters

ParameterTypeDescription
rt_framework_idstrRequired. The id of the red teaming framework.

Returns

This method returns an object of type RTFramework.

Update RT Framework

Renames a red teaming framework or changes its description, and returns it. Its risk categories are managed through their own endpoints.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.rt_frameworks.update(
    rt_framework_id="<RT-FRAMEWORK-ID>",
    name="OWASP Top 10 for LLMs",
    description="Our baseline coverage before each release.",
)

For async mode, call a_update and await it as shown below:

result = await client.rt_frameworks.a_update(...)

Parameters

ParameterTypeDescription
rt_framework_idstrRequired. The id of the red teaming framework.
nameOptional[str]The name of the framework, unique within the project.
descriptionOptional[str]What the framework covers. Send null to clear it.

Returns

This method returns an object of type RTFramework.

Delete RT Framework

Permanently deletes a red teaming framework and its risk categories. Risk assessments already run from it are kept.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.rt_frameworks.delete(rt_framework_id="<RT-FRAMEWORK-ID>")

For async mode, call a_delete and await it as shown below:

result = await client.rt_frameworks.a_delete(...)

Parameters

ParameterTypeDescription
rt_framework_idstrRequired. The id of the red teaming framework.

Returns

This method returns an object of type RTFrameworkRef.

Run RT Framework

Runs the named risk categories of a red teaming framework against your AI connection or prompt, and returns the id of the risk assessment it started. The assessment runs in the background; follow it on the Confident AI platform. Provide exactly one target: aiConnectionId or promptAlias.

from confident_ai import ConfidentAI
from confident_ai.rt_frameworks import AttackEngine
from confident_ai.common import GenerationMode
from confident_ai.common import Level

client = ConfidentAI()

result = client.rt_frameworks.run(
    rt_framework_id="<RT-FRAMEWORK-ID>",
    risk_categories=["Data protection"],
    exposure=Level.LOW,
    identifier="pre-release-2025-01",
    ai_connection_id="<AI-CONNECTION-ID>",
    prompt_alias="customer-support",
    prompt_commit="bab04ce",
    generation_mode=GenerationMode.AI_CONNECTION,
    attack_engine=AttackEngine(
        generation_guidelines=[
            "Write in the voice of a frustrated customer."
        ]
    ),
)

For async mode, call a_run and await it as shown below:

result = await client.rt_frameworks.a_run(...)

Parameters

ParameterTypeDescription
rt_framework_idstrRequired. The id of the red teaming framework.
risk_categoriesList[str]Required. The names of the framework's risk categories to run. Every name must belong to this framework.
exposureLevelRequired. See Level.
identifierOptional[str]A human-readable identifier for the risk assessment, shown on the Confident AI platform.
ai_connection_idOptional[str]The id of the AI connection to attack. Send this or promptAlias, never both.
prompt_aliasOptional[str]The alias of the prompt to attack. Send this or aiConnectionId, never both.
prompt_commitOptional[str]The commit hash of the prompt to attack. Requires promptAlias; defaults to its latest commit.
generation_modeOptional[GenerationMode]See GenerationMode.
attack_engineOptional[AttackEngine]See AttackEngine.

Returns

This method returns an object of type RiskAssessmentRef.

Types

AttackEngine

How attacks are generated for this run.

class AttackEngine:
    generation_guidelines: Optional[List[str]] = Field(default=None, alias="generationGuidelines")

generation_guidelinesOptional[List[str]]

Extra instructions the attacker follows when generating attacks against this system.

Example: ["Write in the voice of a frustrated customer."]

FrameworkAttackMethod

An attack method as a framework carries it, with the settings a run would use.

class FrameworkAttackMethod:
    name: str
    multi_turn: bool = Field(alias="multiTurn")
    parameters: Optional[Dict[str, Any]]

namestrRequired

The name of the attack method.

Example: "Prompt Injection"

multi_turnboolRequired

Whether the attack plays out over a conversation rather than a single request.

Example: false

parametersOptional[Dict[str, Any]]Required

The parameter values this project runs the attack method with, or null when it takes none.

Example: {"persona":"urgent"}

FrameworkRiskCategory

A risk category as the framework carries it, resolved to everything a run needs.

class FrameworkRiskCategory:
    id: str
    name: str
    description: Optional[str]
    vulnerabilities: List[FrameworkVulnerability]
    attack_methods: List[FrameworkAttackMethod] = Field(alias="attackMethods")

idstrRequired

The id of the risk category, generated by Confident AI.

Example: "<RISK-CATEGORY-ID>"

namestrRequired

The name of the risk category, unique within the framework.

Example: "Data protection"

descriptionOptional[str]Required

What this risk category covers.

Example: "Risks around leaking data the model was given."

vulnerabilitiesList[FrameworkVulnerability]Required

The vulnerabilities this category probes for, with their selected types grouped under each one.

See FrameworkVulnerability.

attack_methodsList[FrameworkAttackMethod]Required

The attack methods this category probes with.

See FrameworkAttackMethod.

FrameworkRiskCategorySummary

A framework's risk category, counted rather than resolved.

class FrameworkRiskCategorySummary:
    name: str
    num_vulnerability_types: int = Field(alias="numVulnerabilityTypes")
    num_attack_methods: int = Field(alias="numAttackMethods")

namestrRequired

The name of the risk category.

Example: "Data protection"

num_vulnerability_typesintRequired

How many vulnerability types the category selects.

Example: 4

num_attack_methodsintRequired

How many attack methods the category selects.

Example: 3

FrameworkVulnerability

A vulnerability as a framework carries it: the types this category selected, with everything the evaluator needs to judge a reply.

class FrameworkVulnerability:
    name: str
    types: List[str]
    criteria: Optional[str]
    evaluation_guidelines: List[str] = Field(alias="evaluationGuidelines")
    evaluation_examples: Optional[List[FrameworkVulnerabilityEvaluationExample]] = Field(alias="evaluationExamples")

namestrRequired

The name of the vulnerability.

Example: "Prompt Leakage"

typesList[str]Required

The names of its types this category selects.

Example: ["System prompt disclosure"]

criteriaOptional[str]Required

The rule the evaluator applies to decide whether a reply is vulnerable.

Example: "The output must not reveal the system prompt or its rules."

evaluation_guidelinesList[str]Required

Extra instructions the evaluator follows.

Example: ["Treat a partial quote of the prompt as a failure."]

evaluation_examplesOptional[List[FrameworkVulnerabilityEvaluationExample]]Required

Worked examples that steer the evaluator, or null when the vulnerability has none.

GenerationMode

Where the actual outputs come from when running a dataset: AI_CONNECTION generates them with an AI connection, PROMPT with a prompt. Omit it when you supply at most one of aiConnectionId or promptAlias, and Confident AI infers the mode from whichever you sent.

class GenerationMode(Enum):
    AI_CONNECTION = "AI_CONNECTION"
    PROMPT = "PROMPT"

AI_CONNECTION · PROMPT

Level

A three-step scale, LOW to HIGH, used for how exposed a system under test is and how easily an attack method exploits it.

class Level(Enum):
    LOW = "LOW"
    MEDIUM = "MEDIUM"
    HIGH = "HIGH"

LOW · MEDIUM · HIGH

RTFramework

A red teaming framework: the risk categories a risk assessment runs, resolved to everything the run needs.

class RTFramework:
    id: str
    name: str
    description: Optional[str]
    risk_categories: List[FrameworkRiskCategory] = Field(alias="riskCategories")

idstrRequired

The id of the framework, generated by Confident AI.

Example: "<RT-FRAMEWORK-ID>"

namestrRequired

The name of the framework.

Example: "OWASP Top 10 for LLMs"

descriptionOptional[str]Required

What the framework covers.

Example: "Our baseline coverage before each release."

risk_categoriesList[FrameworkRiskCategory]Required

The framework's risk categories, each resolved to the vulnerabilities and attack methods a run would use.

See FrameworkRiskCategory.

RTFrameworkList

One page of frameworks, with the total across all pages.

class RTFrameworkList:
    rt_frameworks: List[RTFrameworkSummary] = Field(alias="rtFrameworks")
    total_rt_frameworks: int = Field(alias="totalRTFrameworks")
    page: int
    page_size: int = Field(alias="pageSize")

rt_frameworksList[RTFrameworkSummary]Required

The frameworks for the current page, ordered by name.

See RTFrameworkSummary.

total_rt_frameworksintRequired

The total number of frameworks in this project.

Example: 3

pageintRequired

The page this response covers.

Example: 1

page_sizeintRequired

The number of frameworks per page.

Example: 25

RTFrameworkRef

A reference to a red teaming framework by its id.

class RTFrameworkRef:
    id: str

idstrRequired

The id of the framework, generated by Confident AI.

Example: "<RT-FRAMEWORK-ID>"

RTFrameworkSummary

A framework as it appears in a list: its name and what each risk category holds, without resolving the selections.

class RTFrameworkSummary:
    id: str
    name: str
    description: Optional[str]
    risk_categories: List[FrameworkRiskCategorySummary] = Field(alias="riskCategories")

idstrRequired

The id of the framework, generated by Confident AI.

Example: "<RT-FRAMEWORK-ID>"

namestrRequired

The name of the framework.

Example: "OWASP Top 10 for LLMs"

descriptionOptional[str]Required

What the framework covers.

Example: "Our baseline coverage before each release."

risk_categoriesList[FrameworkRiskCategorySummary]Required

The framework's risk categories, counted rather than resolved.

See FrameworkRiskCategorySummary.

RiskAssessmentRef

A reference to a risk assessment by its id.

class RiskAssessmentRef:
    id: str

idstrRequired

The id of the risk assessment the run started.

Example: "<RISK-ASSESSMENT-ID>"

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI