Attack Methods
Overview
The Confident AI SDK exposes every Attack Method method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.
Methods
List Attack Methods
Lists the attack methods available to your Confident AI project one page at a time, ordered by name. The list covers the whole catalog, whether or not this project has configured a method; retrieve one by id for the parameters it takes.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.attack_methods.list(
page=1,
page_size=25,
multi_turn="false",
)For async mode, call a_list and await it as shown below:
result = await client.attack_methods.a_list(...)Parameters
| Parameter | Type | Description |
|---|---|---|
page | Optional[int] | The page of attack methods to return. Defaults to 1. |
page_size | Optional[int] | The number of attack methods per page, at most 100. Defaults to 25. |
multi_turn | Optional[Literal['true', 'false']] | When true, returns only multi-turn attack methods; when false, only single-turn ones. Omit to return both. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.attackMethods.list(
{ page: 1, pageSize: 25, multiTurn: "false" },
);Parameters
| Parameter | Type | Description |
|---|---|---|
page | number | The page of attack methods to return. Defaults to 1. |
pageSize | number | The number of attack methods per page, at most 100. Defaults to 25. |
multiTurn | "true" | "false" | When true, returns only multi-turn attack methods; when false, only single-turn ones. Omit to return both. |
Returns
This method returns an object of type AttackMethodList.
Get Attack Method
Retrieves an attack method by id, with the parameters it takes and the values in force for your project.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.attack_methods.get(attack_method_id="Prompt Injection")For async mode, call a_get and await it as shown below:
result = await client.attack_methods.a_get(...)Parameters
| Parameter | Type | Description |
|---|---|---|
attack_method_id | str | Required. The id of the attack method, as the list returns it. A method's catalog name also resolves. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.attackMethods.get("Prompt Injection");Parameters
| Parameter | Type | Description |
|---|---|---|
attackMethodId | string | Required. The id of the attack method, as the list returns it. A method's catalog name also resolves. |
Returns
This method returns an object of type AttackMethod.
Update Attack Method
Configures the parameter values your project runs an attack method with, and returns the method. An attack method that takes no parameters cannot be configured.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.attack_methods.update(
attack_method_id="Prompt Injection",
parameters={"persona": "urgent"},
)For async mode, call a_update and await it as shown below:
result = await client.attack_methods.a_update(...)Parameters
| Parameter | Type | Description |
|---|---|---|
attack_method_id | str | Required. The id of the attack method, as the list returns it. A method's catalog name also resolves. |
parameters | Dict[str, Any] | Required. The values to configure, keyed by parameter name. They replace this project's stored configuration wholesale, so send every value you want kept, and each must match the type its parameter declares. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.attackMethods.update(
"Prompt Injection",
{ persona: "urgent" },
);Parameters
| Parameter | Type | Description |
|---|---|---|
attackMethodId | string | Required. The id of the attack method, as the list returns it. A method's catalog name also resolves. |
parameters | Record<string, unknown> | Required. The values to configure, keyed by parameter name. They replace this project's stored configuration wholesale, so send every value you want kept, and each must match the type its parameter declares. |
Returns
This method returns an object of type AttackMethod.
Reset Attack Method
Clears your project's configuration of an attack method, so it runs with the catalog defaults again. The attack method itself belongs to the Confident AI catalog and is not deleted, which makes this safe to repeat.
from confident_ai import ConfidentAI
client = ConfidentAI()
result = client.attack_methods.reset(attack_method_id="Prompt Injection")For async mode, call a_reset and await it as shown below:
result = await client.attack_methods.a_reset(...)Parameters
| Parameter | Type | Description |
|---|---|---|
attack_method_id | str | Required. The id of the attack method, as the list returns it. A method's catalog name also resolves. |
import { ConfidentAI } from "confident-ai";
const client = new ConfidentAI();
const result = await client.attackMethods.reset("Prompt Injection");Parameters
| Parameter | Type | Description |
|---|---|---|
attackMethodId | string | Required. The id of the attack method, as the list returns it. A method's catalog name also resolves. |
Returns
This method returns an object of type AttackMethodRef.
Types
AttackMethod
One way of attacking the system under test, drawn from the Confident AI catalog and configured per project.
class AttackMethod:
id: str
name: str
description: Optional[str]
multi_turn: bool = Field(alias="multiTurn")
exploitability: Optional[Level]
configurable: bool
parameters: Optional[Dict[str, AttackParameter]]idstrRequired
The id of the attack method. It is the method's catalog name until this project configures it, and its generated id afterwards; both keep resolving.
Example: "Prompt Injection"
namestrRequired
The name of the attack method, as the catalog spells it.
Example: "Prompt Injection"
descriptionOptional[str]Required
What the attack method does to the system under test.
Example: "Embeds instructions in the input that try to override the system prompt."
multi_turnboolRequired
Whether the attack plays out over a conversation rather than a single request.
Example: false
exploitabilityOptional[Level]Required
How easily the attack method exploits a vulnerable system.
See Level.
configurableboolRequired
Whether the attack method takes parameters. Updating one that takes none is rejected.
Example: true
parametersOptional[Dict[str, AttackParameter]]Required
The parameters the attack method takes, keyed by parameter name, each carrying the value in force for this project. Null when the method takes none.
See AttackParameter.
interface AttackMethod {
id: string;
name: string;
description: string | null;
multiTurn: boolean;
exploitability: Level | null;
configurable: boolean;
parameters: Record<string, AttackParameter> | null;
}idstringRequired
The id of the attack method. It is the method's catalog name until this project configures it, and its generated id afterwards; both keep resolving.
Example: "Prompt Injection"
namestringRequired
The name of the attack method, as the catalog spells it.
Example: "Prompt Injection"
descriptionstring | nullRequired
What the attack method does to the system under test.
Example: "Embeds instructions in the input that try to override the system prompt."
multiTurnbooleanRequired
Whether the attack plays out over a conversation rather than a single request.
Example: false
exploitabilityLevel | nullRequired
How easily the attack method exploits a vulnerable system.
See Level.
configurablebooleanRequired
Whether the attack method takes parameters. Updating one that takes none is rejected.
Example: true
parametersRecord<string, AttackParameter> | nullRequired
The parameters the attack method takes, keyed by parameter name, each carrying the value in force for this project. Null when the method takes none.
See AttackParameter.
AttackMethodList
One page of attack methods, with the total across all pages.
class AttackMethodList:
attack_methods: List[AttackMethodSummary] = Field(alias="attackMethods")
total_attack_methods: int = Field(alias="totalAttackMethods")
page: int
page_size: int = Field(alias="pageSize")attack_methodsList[AttackMethodSummary]Required
The attack methods for the current page, ordered by name.
See AttackMethodSummary.
total_attack_methodsintRequired
The total number of attack methods matching the query across all pages.
Example: 42
pageintRequired
The page this response covers.
Example: 1
page_sizeintRequired
The number of attack methods per page.
Example: 25
interface AttackMethodList {
attackMethods: AttackMethodSummary[];
totalAttackMethods: number;
page: number;
pageSize: number;
}attackMethodsAttackMethodSummary[]Required
The attack methods for the current page, ordered by name.
See AttackMethodSummary.
totalAttackMethodsnumberRequired
The total number of attack methods matching the query across all pages.
Example: 42
pagenumberRequired
The page this response covers.
Example: 1
pageSizenumberRequired
The number of attack methods per page.
Example: 25
AttackMethodRef
A reference to an attack method by its id.
class AttackMethodRef:
id: stridstrRequired
The id of the attack method.
Example: "Prompt Injection"
interface AttackMethodRef {
id: string;
}idstringRequired
The id of the attack method.
Example: "Prompt Injection"
AttackMethodSummary
An attack method as it appears in a list, without its parameters.
class AttackMethodSummary:
id: str
name: str
description: Optional[str]
multi_turn: bool = Field(alias="multiTurn")
exploitability: Optional[Level]
configurable: boolidstrRequired
The id of the attack method. It is the method's catalog name until this project configures it, and its generated id afterwards; both keep resolving.
Example: "Prompt Injection"
namestrRequired
The name of the attack method, as the catalog spells it.
Example: "Prompt Injection"
descriptionOptional[str]Required
What the attack method does to the system under test.
Example: "Embeds instructions in the input that try to override the system prompt."
multi_turnboolRequired
Whether the attack plays out over a conversation rather than a single request.
Example: false
exploitabilityOptional[Level]Required
How easily the attack method exploits a vulnerable system.
See Level.
configurableboolRequired
Whether the attack method takes parameters. Updating one that takes none is rejected.
Example: true
interface AttackMethodSummary {
id: string;
name: string;
description: string | null;
multiTurn: boolean;
exploitability: Level | null;
configurable: boolean;
}idstringRequired
The id of the attack method. It is the method's catalog name until this project configures it, and its generated id afterwards; both keep resolving.
Example: "Prompt Injection"
namestringRequired
The name of the attack method, as the catalog spells it.
Example: "Prompt Injection"
descriptionstring | nullRequired
What the attack method does to the system under test.
Example: "Embeds instructions in the input that try to override the system prompt."
multiTurnbooleanRequired
Whether the attack plays out over a conversation rather than a single request.
Example: false
exploitabilityLevel | nullRequired
How easily the attack method exploits a vulnerable system.
See Level.
configurablebooleanRequired
Whether the attack method takes parameters. Updating one that takes none is rejected.
Example: true
AttackParameter
One setting an attack method takes, with the value in force for this project.
class AttackParameter:
type: AttackParameterType
required: Optional[bool] = None
description: Optional[str] = None
options: Optional[List[str]] = None
default: Optional[Any] = NonetypeAttackParameterTypeRequired
See AttackParameterType.
requiredOptional[bool]
Whether the attack method refuses to run without this parameter.
Example: true
descriptionOptional[str]
What the parameter controls.
Example: "The persona the attacker adopts."
optionsOptional[List[str]]
The values an enum parameter accepts.
Example: ["polite","urgent"]
defaultOptional[Any]
The value this project has configured, or the catalog default when it has configured none.
Example: "urgent"
interface AttackParameter {
type: AttackParameterType;
required?: boolean;
description?: string;
options?: string[];
default?: unknown;
}typeAttackParameterTypeRequired
See AttackParameterType.
requiredboolean
Whether the attack method refuses to run without this parameter.
Example: true
descriptionstring
What the parameter controls.
Example: "The persona the attacker adopts."
optionsstring[]
The values an enum parameter accepts.
Example: ["polite","urgent"]
defaultunknown
The value this project has configured, or the catalog default when it has configured none.
Example: "urgent"
AttackParameterType
The kind of value a parameter takes. enum accepts one of the strings in options, and json accepts an object.
class AttackParameterType(Enum):
STRING = "string"
INTEGER = "integer"
FLOAT = "float"
BOOLEAN = "boolean"
ENUM = "enum"
JSON = "json"enum AttackParameterType {
STRING = "string",
INTEGER = "integer",
FLOAT = "float",
BOOLEAN = "boolean",
ENUM = "enum",
JSON = "json",
}STRING · INTEGER · FLOAT · BOOLEAN · ENUM · JSON
Level
A three-step scale, LOW to HIGH, used for how exposed a system under test is and how easily an attack method exploits it.
class Level(Enum):
LOW = "LOW"
MEDIUM = "MEDIUM"
HIGH = "HIGH"enum Level {
LOW = "LOW",
MEDIUM = "MEDIUM",
HIGH = "HIGH",
}LOW · MEDIUM · HIGH
Last updated on