Launch Week 3: Five days of launches

Attack Methods

Overview

The Confident AI SDK exposes every Attack Method method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Attack Methods

Lists the attack methods available to your Confident AI project one page at a time, ordered by name. The list covers the whole catalog, whether or not this project has configured a method; retrieve one by id for the parameters it takes.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.attack_methods.list(
    page=1,
    page_size=25,
    multi_turn="false",
)

For async mode, call a_list and await it as shown below:

result = await client.attack_methods.a_list(...)

Parameters

ParameterTypeDescription
pageOptional[int]The page of attack methods to return. Defaults to 1.
page_sizeOptional[int]The number of attack methods per page, at most 100. Defaults to 25.
multi_turnOptional[Literal['true', 'false']]When true, returns only multi-turn attack methods; when false, only single-turn ones. Omit to return both.

Returns

This method returns an object of type AttackMethodList.

Get Attack Method

Retrieves an attack method by id, with the parameters it takes and the values in force for your project.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.attack_methods.get(attack_method_id="Prompt Injection")

For async mode, call a_get and await it as shown below:

result = await client.attack_methods.a_get(...)

Parameters

ParameterTypeDescription
attack_method_idstrRequired. The id of the attack method, as the list returns it. A method's catalog name also resolves.

Returns

This method returns an object of type AttackMethod.

Update Attack Method

Configures the parameter values your project runs an attack method with, and returns the method. An attack method that takes no parameters cannot be configured.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.attack_methods.update(
    attack_method_id="Prompt Injection",
    parameters={"persona": "urgent"},
)

For async mode, call a_update and await it as shown below:

result = await client.attack_methods.a_update(...)

Parameters

ParameterTypeDescription
attack_method_idstrRequired. The id of the attack method, as the list returns it. A method's catalog name also resolves.
parametersDict[str, Any]Required. The values to configure, keyed by parameter name. They replace this project's stored configuration wholesale, so send every value you want kept, and each must match the type its parameter declares.

Returns

This method returns an object of type AttackMethod.

Reset Attack Method

Clears your project's configuration of an attack method, so it runs with the catalog defaults again. The attack method itself belongs to the Confident AI catalog and is not deleted, which makes this safe to repeat.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.attack_methods.reset(attack_method_id="Prompt Injection")

For async mode, call a_reset and await it as shown below:

result = await client.attack_methods.a_reset(...)

Parameters

ParameterTypeDescription
attack_method_idstrRequired. The id of the attack method, as the list returns it. A method's catalog name also resolves.

Returns

This method returns an object of type AttackMethodRef.

Types

AttackMethod

One way of attacking the system under test, drawn from the Confident AI catalog and configured per project.

class AttackMethod:
    id: str
    name: str
    description: Optional[str]
    multi_turn: bool = Field(alias="multiTurn")
    exploitability: Optional[Level]
    configurable: bool
    parameters: Optional[Dict[str, AttackParameter]]

idstrRequired

The id of the attack method. It is the method's catalog name until this project configures it, and its generated id afterwards; both keep resolving.

Example: "Prompt Injection"

namestrRequired

The name of the attack method, as the catalog spells it.

Example: "Prompt Injection"

descriptionOptional[str]Required

What the attack method does to the system under test.

Example: "Embeds instructions in the input that try to override the system prompt."

multi_turnboolRequired

Whether the attack plays out over a conversation rather than a single request.

Example: false

exploitabilityOptional[Level]Required

How easily the attack method exploits a vulnerable system.

See Level.

configurableboolRequired

Whether the attack method takes parameters. Updating one that takes none is rejected.

Example: true

parametersOptional[Dict[str, AttackParameter]]Required

The parameters the attack method takes, keyed by parameter name, each carrying the value in force for this project. Null when the method takes none.

See AttackParameter.

AttackMethodList

One page of attack methods, with the total across all pages.

class AttackMethodList:
    attack_methods: List[AttackMethodSummary] = Field(alias="attackMethods")
    total_attack_methods: int = Field(alias="totalAttackMethods")
    page: int
    page_size: int = Field(alias="pageSize")

attack_methodsList[AttackMethodSummary]Required

The attack methods for the current page, ordered by name.

See AttackMethodSummary.

total_attack_methodsintRequired

The total number of attack methods matching the query across all pages.

Example: 42

pageintRequired

The page this response covers.

Example: 1

page_sizeintRequired

The number of attack methods per page.

Example: 25

AttackMethodRef

A reference to an attack method by its id.

class AttackMethodRef:
    id: str

idstrRequired

The id of the attack method.

Example: "Prompt Injection"

AttackMethodSummary

An attack method as it appears in a list, without its parameters.

class AttackMethodSummary:
    id: str
    name: str
    description: Optional[str]
    multi_turn: bool = Field(alias="multiTurn")
    exploitability: Optional[Level]
    configurable: bool

idstrRequired

The id of the attack method. It is the method's catalog name until this project configures it, and its generated id afterwards; both keep resolving.

Example: "Prompt Injection"

namestrRequired

The name of the attack method, as the catalog spells it.

Example: "Prompt Injection"

descriptionOptional[str]Required

What the attack method does to the system under test.

Example: "Embeds instructions in the input that try to override the system prompt."

multi_turnboolRequired

Whether the attack plays out over a conversation rather than a single request.

Example: false

exploitabilityOptional[Level]Required

How easily the attack method exploits a vulnerable system.

See Level.

configurableboolRequired

Whether the attack method takes parameters. Updating one that takes none is rejected.

Example: true

AttackParameter

One setting an attack method takes, with the value in force for this project.

class AttackParameter:
    type: AttackParameterType
    required: Optional[bool] = None
    description: Optional[str] = None
    options: Optional[List[str]] = None
    default: Optional[Any] = None

typeAttackParameterTypeRequired

requiredOptional[bool]

Whether the attack method refuses to run without this parameter.

Example: true

descriptionOptional[str]

What the parameter controls.

Example: "The persona the attacker adopts."

optionsOptional[List[str]]

The values an enum parameter accepts.

Example: ["polite","urgent"]

defaultOptional[Any]

The value this project has configured, or the catalog default when it has configured none.

Example: "urgent"

AttackParameterType

The kind of value a parameter takes. enum accepts one of the strings in options, and json accepts an object.

class AttackParameterType(Enum):
    STRING = "string"
    INTEGER = "integer"
    FLOAT = "float"
    BOOLEAN = "boolean"
    ENUM = "enum"
    JSON = "json"

STRING · INTEGER · FLOAT · BOOLEAN · ENUM · JSON

Level

A three-step scale, LOW to HIGH, used for how exposed a system under test is and how easily an attack method exploits it.

class Level(Enum):
    LOW = "LOW"
    MEDIUM = "MEDIUM"
    HIGH = "HIGH"

LOW · MEDIUM · HIGH

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI