Launch Week 02 wrapped — explore all five launches

Create Rule

POSThttps://api.confident-ai.com/v2/evaluation-rules

Creates a standing rule that runs a metric collection against matching production traces, spans or threads as they arrive, and returns its id. Running metrics consumes LLM usage. The metric collection's turn type must match the rule: THREAD rules require a multi-turn collection, TRACE and SPAN rules a single-turn one, and only one enabled THREAD rule may target a given collection. Requires the Starter plan or above.

POST/v2/evaluation-rules
curl -X POST "https://api.confident-ai.com/v2/evaluation-rules" \
  -H "CONFIDENT_API_KEY: <PROJECT-API-KEY>" \
  -H "Content-Type: application/json" \
  -d '{
  "name": "Score production answers",
  "enabled": true,
  "dataModel": "TRACE",
  "metricCollectionId": "<METRIC-COLLECTION-ID>",
  "description": "Scores answers we serve to end users.",
  "sampleRate": 0.2,
  "spanType": "SPAN",
  "filters": {
    "operator": "AND",
    "groups": [
      {
        "operator": "AND",
        "filters": [
          {
            "category": "Name",
            "condition": "Is",
            "value": "capital-lookup"
          }
        ]
      }
    ]
  },
  "threadTimelimit": 300,
  "overwriteEvals": false
}'
200
{
  "success": true,
  "data": {
    "id": "<EVALUATION-RULE-ID>"
  },
  "link": "https://app.confident-ai.com/project/<PROJECT-ID>/workflows",
  "deprecated": false
}

Headers

  • CONFIDENT_API_KEYstringRequired

    The API key of your Confident AI project.

Request body

  • namestringRequired

    A name for the rule, unique within the project.

  • enabledboolean

    Whether the rule evaluates matching items as they arrive. Defaults to true.

  • dataModelenumRequired

    What kind of production item a rule evaluates: TRACE for a whole trace, SPAN for a single step within one, THREAD for a finished conversation.

    Show 3 enum valuesHide 3 enum values
    • TRACE
    • SPAN
    • THREAD
  • metricCollectionIdstringRequired

    The id of the metric collection to run. It must be multi-turn for THREAD rules and single-turn for TRACE and SPAN rules.

  • descriptionstring | null

    A note about what the rule checks. Send null to clear it.

  • sampleRatenumber

    The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them.

  • spanTypeenum | null

    The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.

    Show 5 enum valuesHide 5 enum values
    • SPAN
    • AGENT
    • TOOL
    • RETRIEVER
    • LLM
  • filtersobject | null

    A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.

    Show 2 propertiesHide 2 properties
    • operatorenumRequired

      Show 2 enum valuesHide 2 enum values
      • AND
      • OR
    • groupslist of objectsRequired

      Show 2 propertiesHide 2 properties
      • operatorenumRequired

        Show 2 enum valuesHide 2 enum values
        • AND
        • OR
      • filterslist of objectsRequired

        Show 4 propertiesHide 4 properties
        • categoryenumRequired

          Show 80 enum valuesHide 80 enum values
          • User Id
          • Thread Id
          • Trace Uuid
          • Trace Name
          • Trace Version
          • Trace Status
          • Trace Tags
          • Trace
          • Span Uuid
          • Name
          • Span Name
          • Span Type
          • Span Status
          • Metrics Status
          • Error Status
          • Name
          • Model
          • Provider
          • Integration
          • Embedder
          • Chunk Size
          • Top-K
          • Hyperparameter
          • Dataset
          • Dataset Name
          • Test Run ID
          • Identifier
          • Test File
          • Status
          • Official
          • Evals Mode
          • Tests Passed
          • Tests Failed
          • Pass Rate
          • Fail Rate
          • Star Rating
          • Thumbs Rating
          • Explanation
          • Expected Output
          • Expected Outcome
          • Annotator
          • End User
          • Annotation Type
          • Annotation Name
          • Criteria
          • Annotation Date
          • Metric Score
          • Metric Status
          • Name
          • Metadata
          • Classifier
          • Metric
          • Metric Name
          • Trace Count
          • Test Case ID
          • Requested review from
          • Assigned to
          • Tags
          • Labels
          • Tools Called
          • Finalized
          • Golden ID
          • Ingestion Task
          • Latency
          • Environment
          • Review flag
          • Vulnerability
          • Vulnerability Type
          • Attack Method
          • Risk Category
          • Framework
          • Assessment ID
          • Prompt Alias
          • Prompt Version
          • Prompt Label
          • Prompt Commit Hash
          • Prompt
          • Annotations
          • Status Code
          • Actor Type
        • conditionenum | enum | enum | enum | enum | enum | enum | enum | enum | enumRequired

          Show 10 variantsHide 10 variants
          • enum

            Show 6 enum valuesHide 6 enum values
            • Is less than
            • Is equal or less than
            • Is greater than
            • Is equal or greater than
            • Is equal to
            • Does not equal
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Has
            • Has not
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Is
            • Is not
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Is one of
            • Is not one of
          • OR
          • enum

            Show 4 enum valuesHide 4 enum values
            • Is
            • Is not
            • Is empty
            • Is not empty
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Contains
            • Does not contain
          • OR
          • enum

            Show 3 enum valuesHide 3 enum values
            • Contains
            • Contains only
            • Does not contain
          • OR
          • enum

            Show 4 enum valuesHide 4 enum values
            • Has decreased by more than
            • Has decreased by less than
            • Has increased by more than
            • Has increased by less than
          • OR
          • enum

            Show 1 enum valueHide 1 enum value
            • Has changed from
          • OR
          • enum

            Show 1 enum valueHide 1 enum value
            • Is between
        • valuestring | number | list of stringsRequired

          Show 3 variantsHide 3 variants
          • string

          • OR
          • number

          • OR
          • list of strings

        • keystring

  • threadTimelimitinteger

    For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. Defaults to 300.

  • overwriteEvalsboolean

    Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false.

Response

Create Rule succeeded.

  • successboolean

    Indicates if the request was successful.

  • dataobject

    A reference to an evaluation rule by its id.

    Show 1 propertyHide 1 property
    • idstring

      The id of the rule, generated by Confident AI.

  • linkstring

    This is the URL of the resource on the Confident AI platform.

  • deprecatedboolean

    Indicates if this endpoint is deprecated.

Built byConfident AI