Launch Week 3: Five days of launches

Create Evaluation Rule

(v2)

POST

Creates a standing rule that runs a metric collection against matching production traces, spans or threads as they arrive, and returns its id. Running metrics consumes LLM usage. The metric collection's turn type must match the rule: THREAD rules require a multi-turn collection, TRACE and SPAN rules a single-turn one, and only one enabled THREAD rule may target a given collection. Requires the Starter plan or above.

POST/v2/evaluation-rules
curl -X POST "https://api.confident-ai.com/v2/evaluation-rules" \
  -H "CONFIDENT_API_KEY: <PROJECT-API-KEY>" \
  -H "Content-Type: application/json" \
  -d '{
  "name": "Score production answers",
  "enabled": true,
  "dataModel": "TRACE",
  "metricCollectionId": "<METRIC-COLLECTION-ID>",
  "description": "Scores answers we serve to end users.",
  "sampleRate": 0.2,
  "spanType": "SPAN",
  "filters": {
    "operator": "AND",
    "groups": [
      {
        "operator": "AND",
        "filters": [
          {
            "category": "Name",
            "condition": "Is",
            "value": "capital-lookup"
          }
        ]
      }
    ]
  },
  "threadTimelimit": 300,
  "overwriteEvals": false
}'
200
{
  "success": true,
  "data": {
    "id": "<EVALUATION-RULE-ID>"
  },
  "link": "https://app.confident-ai.com/project/<PROJECT-ID>/workflows",
  "deprecated": false
}

Headers

  • CONFIDENT_API_KEYstringRequired

    The API key of your Confident AI project.

Request body

  • namestringRequired

    A name for the rule, unique within the project.

  • enabledboolean

    Whether the rule evaluates matching items as they arrive. Defaults to true.

  • dataModelenumRequired

    What kind of production item a rule evaluates: TRACE for a whole trace, SPAN for a single step within one, THREAD for a finished conversation.

    Show 3 enum valuesHide 3 enum values
    • TRACE
    • SPAN
    • THREAD
  • metricCollectionIdstringRequired

    The id of the metric collection to run. It must be multi-turn for THREAD rules and single-turn for TRACE and SPAN rules.

  • descriptionstring | null

    A note about what the rule checks. Send null to clear it.

  • sampleRatenumber

    The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them.

  • spanTypeenum | null

    The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.

    Show 5 enum valuesHide 5 enum values
    • SPAN
    • AGENT
    • TOOL
    • RETRIEVER
    • LLM
  • filtersobject | null

    A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.

    Show 2 propertiesHide 2 properties
    • operatorenumRequired

      Show 2 enum valuesHide 2 enum values
      • AND
      • OR
    • groupslist of objectsRequired

      Show 2 propertiesHide 2 properties
      • operatorenumRequired

        Show 2 enum valuesHide 2 enum values
        • AND
        • OR
      • filterslist of objectsRequired

        Show 4 propertiesHide 4 properties
        • categoryenumRequired

          Show 80 enum valuesHide 80 enum values
          • User Id
          • Thread Id
          • Trace UUID
          • Trace Name
          • Trace Version
          • Trace Status
          • Trace Tags
          • Trace
          • Span UUID
          • Name
          • Span Name
          • Span Type
          • Span Status
          • Metrics Status
          • Error Status
          • Name
          • Model
          • Provider
          • Integration
          • Embedder
          • Chunk Size
          • Top-K
          • Hyperparameter
          • Dataset
          • Dataset Name
          • Test Run ID
          • Identifier
          • Test File
          • Status
          • Official
          • Evals Mode
          • Tests Passed
          • Tests Failed
          • Pass Rate
          • Fail Rate
          • Star Rating
          • Thumbs Rating
          • Explanation
          • Expected Output
          • Expected Outcome
          • Annotator
          • End User
          • Annotation Type
          • Annotation Name
          • Criteria
          • Annotation Date
          • Metric Score
          • Metric Status
          • Name
          • Metadata
          • Classifier
          • Metric
          • Metric Name
          • Trace Count
          • Test Case ID
          • Requested review from
          • Assigned to
          • Tags
          • Labels
          • Tools Called
          • Finalized
          • Golden ID
          • Ingestion Task
          • Latency
          • Environment
          • Review flag
          • Vulnerability
          • Vulnerability Type
          • Attack Method
          • Risk Category
          • Framework
          • Assessment ID
          • Prompt Alias
          • Prompt Version
          • Prompt Label
          • Prompt Commit Hash
          • Prompt
          • Annotations
          • Status Code
          • Actor Type
        • conditionenum | enum | enum | enum | enum | enum | enum | enum | enum | enumRequired

          Show 10 variantsHide 10 variants
          • enum

            Show 6 enum valuesHide 6 enum values
            • Is less than
            • Is equal or less than
            • Is greater than
            • Is equal or greater than
            • Is equal to
            • Does not equal
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Has
            • Has not
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Is
            • Is not
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Is one of
            • Is not one of
          • OR
          • enum

            Show 4 enum valuesHide 4 enum values
            • Is
            • Is not
            • Is empty
            • Is not empty
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Contains
            • Does not contain
          • OR
          • enum

            Show 3 enum valuesHide 3 enum values
            • Contains
            • Contains only
            • Does not contain
          • OR
          • enum

            Show 4 enum valuesHide 4 enum values
            • Has decreased by more than
            • Has decreased by less than
            • Has increased by more than
            • Has increased by less than
          • OR
          • enum

            Show 1 enum valueHide 1 enum value
            • Has changed from
          • OR
          • enum

            Show 1 enum valueHide 1 enum value
            • Is between
        • valuestring | number | list of stringsRequired

          Show 3 variantsHide 3 variants
          • string

          • OR
          • number

          • OR
          • list of strings

        • keystring

  • threadTimelimitinteger | null

    For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300.

  • overwriteEvalsboolean

    Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false.

Response

Create Evaluation Rule succeeded.

  • successboolean

    Indicates if the request was successful.

  • dataobject

    A reference to an evaluation rule by its id.

    Show 1 propertyHide 1 property
    • idstring

      The id of the rule, generated by Confident AI.

  • linkstring

    This is the URL of the resource on the Confident AI platform.

  • deprecatedboolean

    Indicates if this endpoint is deprecated.

Built byConfident AI