Launch Week 3: Five days of launches

Update Evaluation Rule

(v2)

PUT

Updates an evaluation rule and returns it. Only the fields you send are changed; omitting a field leaves it untouched, and sending null clears it. Constraints are re-checked against the rule the update produces, not just the fields you sent, so switching a rule to THREAD still requires a multi-turn metric collection.

PUT/v2/evaluation-rules/{evaluationRuleId}
curl -X PUT "https://api.confident-ai.com/v2/evaluation-rules/{evaluationRuleId}" \
  -H "CONFIDENT_API_KEY: <PROJECT-API-KEY>" \
  -H "Content-Type: application/json" \
  -d '{
  "name": "Score production answers",
  "enabled": false,
  "dataModel": "TRACE",
  "metricCollectionId": "<METRIC-COLLECTION-ID>",
  "description": "Scores answers we serve to end users.",
  "sampleRate": 0.2,
  "spanType": "SPAN",
  "filters": {
    "operator": "AND",
    "groups": [
      {
        "operator": "AND",
        "filters": [
          {
            "category": "Name",
            "condition": "Is",
            "value": "capital-lookup"
          }
        ]
      }
    ]
  },
  "threadTimelimit": 300,
  "overwriteEvals": false
}'
200
{
  "success": true,
  "data": {
    "id": "<EVALUATION-RULE-ID>",
    "name": "Score production answers",
    "description": "Scores answers we serve to end users.",
    "enabled": true,
    "sampleRate": 0.2,
    "dataModel": "TRACE",
    "spanType": "SPAN",
    "filters": {
      "operator": "AND",
      "groups": [
        {
          "operator": "AND",
          "filters": [
            {
              "category": "Name",
              "condition": "Is",
              "value": "capital-lookup"
            }
          ]
        }
      ]
    },
    "threadTimelimit": 300,
    "overwriteEvals": false,
    "metricCollectionId": "<METRIC-COLLECTION-ID>",
    "createdAt": "2025-01-15T10:30:00.000Z",
    "updatedAt": "2025-01-20T08:00:00.000Z"
  },
  "link": "https://app.confident-ai.com/project/<PROJECT-ID>/workflows",
  "deprecated": false
}

Headers

  • CONFIDENT_API_KEYstringRequired

    The API key of your Confident AI project.

Path parameters

  • evaluationRuleIdstringRequired

    The id of the evaluation rule.

Request body

  • namestring

    A new name for the rule, unique within the project.

  • enabledboolean

    Whether the rule evaluates matching items as they arrive.

  • dataModelenum

    What kind of production item a rule evaluates: TRACE for a whole trace, SPAN for a single step within one, THREAD for a finished conversation.

    Show 3 enum valuesHide 3 enum values
    • TRACE
    • SPAN
    • THREAD
  • metricCollectionIdstring

    The id of a different metric collection to run.

  • descriptionstring | null

    A note about what the rule checks. Send null to clear it.

  • sampleRatenumber

    The fraction of matching items to evaluate, between 0 and 1. Defaults to 1, all of them.

  • spanTypeenum | null

    The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.

    Show 5 enum valuesHide 5 enum values
    • SPAN
    • AGENT
    • TOOL
    • RETRIEVER
    • LLM
  • filtersobject | null

    A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.

    Show 2 propertiesHide 2 properties
    • operatorenumRequired

      Show 2 enum valuesHide 2 enum values
      • AND
      • OR
    • groupslist of objectsRequired

      Show 2 propertiesHide 2 properties
      • operatorenumRequired

        Show 2 enum valuesHide 2 enum values
        • AND
        • OR
      • filterslist of objectsRequired

        Show 4 propertiesHide 4 properties
        • categoryenumRequired

          Show 80 enum valuesHide 80 enum values
          • User Id
          • Thread Id
          • Trace UUID
          • Trace Name
          • Trace Version
          • Trace Status
          • Trace Tags
          • Trace
          • Span UUID
          • Name
          • Span Name
          • Span Type
          • Span Status
          • Metrics Status
          • Error Status
          • Name
          • Model
          • Provider
          • Integration
          • Embedder
          • Chunk Size
          • Top-K
          • Hyperparameter
          • Dataset
          • Dataset Name
          • Test Run ID
          • Identifier
          • Test File
          • Status
          • Official
          • Evals Mode
          • Tests Passed
          • Tests Failed
          • Pass Rate
          • Fail Rate
          • Star Rating
          • Thumbs Rating
          • Explanation
          • Expected Output
          • Expected Outcome
          • Annotator
          • End User
          • Annotation Type
          • Annotation Name
          • Criteria
          • Annotation Date
          • Metric Score
          • Metric Status
          • Name
          • Metadata
          • Classifier
          • Metric
          • Metric Name
          • Trace Count
          • Test Case ID
          • Requested review from
          • Assigned to
          • Tags
          • Labels
          • Tools Called
          • Finalized
          • Golden ID
          • Ingestion Task
          • Latency
          • Environment
          • Review flag
          • Vulnerability
          • Vulnerability Type
          • Attack Method
          • Risk Category
          • Framework
          • Assessment ID
          • Prompt Alias
          • Prompt Version
          • Prompt Label
          • Prompt Commit Hash
          • Prompt
          • Annotations
          • Status Code
          • Actor Type
        • conditionenum | enum | enum | enum | enum | enum | enum | enum | enum | enumRequired

          Show 10 variantsHide 10 variants
          • enum

            Show 6 enum valuesHide 6 enum values
            • Is less than
            • Is equal or less than
            • Is greater than
            • Is equal or greater than
            • Is equal to
            • Does not equal
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Has
            • Has not
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Is
            • Is not
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Is one of
            • Is not one of
          • OR
          • enum

            Show 4 enum valuesHide 4 enum values
            • Is
            • Is not
            • Is empty
            • Is not empty
          • OR
          • enum

            Show 2 enum valuesHide 2 enum values
            • Contains
            • Does not contain
          • OR
          • enum

            Show 3 enum valuesHide 3 enum values
            • Contains
            • Contains only
            • Does not contain
          • OR
          • enum

            Show 4 enum valuesHide 4 enum values
            • Has decreased by more than
            • Has decreased by less than
            • Has increased by more than
            • Has increased by less than
          • OR
          • enum

            Show 1 enum valueHide 1 enum value
            • Has changed from
          • OR
          • enum

            Show 1 enum valueHide 1 enum value
            • Is between
        • valuestring | number | list of stringsRequired

          Show 3 variantsHide 3 variants
          • string

          • OR
          • number

          • OR
          • list of strings

        • keystring

  • threadTimelimitinteger | null

    For THREAD rules, the seconds of inactivity to wait before evaluating a thread, so an in-progress conversation is not scored halfway. The minimum is 120, which leaves time for the last traces to be stored. Send null to use the project's thread timelimit, which defaults to 300.

  • overwriteEvalsboolean

    Re-evaluate items that already have results for this metric collection instead of skipping them. Defaults to false.

Response

Update Evaluation Rule succeeded.

  • successboolean

    Indicates if the request was successful.

  • dataobject

    A standing rule that runs a metric collection against matching production traces, spans or threads as they arrive.

    Show 13 propertiesHide 13 properties
    • idstring

      The id of the rule, generated by Confident AI.

    • namestring

      The name of the rule.

    • descriptionstring | null

      A note about what the rule checks, or null when unset.

    • enabledboolean

      Whether the rule is currently evaluating.

    • sampleRatenumber

      The fraction of matching items the rule evaluates, between 0 and 1.

    • dataModelenum

      What kind of production item a rule evaluates: TRACE for a whole trace, SPAN for a single step within one, THREAD for a finished conversation.

      Show 3 enum valuesHide 3 enum values
      • TRACE
      • SPAN
      • THREAD
    • spanTypeenum | null

      The kind of work a span records: SPAN for a plain step, LLM for a model call, RETRIEVER for a knowledge-base lookup, TOOL for a tool call, and AGENT for an agent step.

      Show 5 enum valuesHide 5 enum values
      • SPAN
      • AGENT
      • TOOL
      • RETRIEVER
      • LLM
    • filtersobject | null

      A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.

      Show 2 propertiesHide 2 properties
      • operatorenum

        Show 2 enum valuesHide 2 enum values
        • AND
        • OR
      • groupslist of objects

        Show 2 propertiesHide 2 properties
        • operatorenum

          Show 2 enum valuesHide 2 enum values
          • AND
          • OR
        • filterslist of objects

          Show 4 propertiesHide 4 properties
          • categoryenum

            Show 80 enum valuesHide 80 enum values
            • User Id
            • Thread Id
            • Trace UUID
            • Trace Name
            • Trace Version
            • Trace Status
            • Trace Tags
            • Trace
            • Span UUID
            • Name
            • Span Name
            • Span Type
            • Span Status
            • Metrics Status
            • Error Status
            • Name
            • Model
            • Provider
            • Integration
            • Embedder
            • Chunk Size
            • Top-K
            • Hyperparameter
            • Dataset
            • Dataset Name
            • Test Run ID
            • Identifier
            • Test File
            • Status
            • Official
            • Evals Mode
            • Tests Passed
            • Tests Failed
            • Pass Rate
            • Fail Rate
            • Star Rating
            • Thumbs Rating
            • Explanation
            • Expected Output
            • Expected Outcome
            • Annotator
            • End User
            • Annotation Type
            • Annotation Name
            • Criteria
            • Annotation Date
            • Metric Score
            • Metric Status
            • Name
            • Metadata
            • Classifier
            • Metric
            • Metric Name
            • Trace Count
            • Test Case ID
            • Requested review from
            • Assigned to
            • Tags
            • Labels
            • Tools Called
            • Finalized
            • Golden ID
            • Ingestion Task
            • Latency
            • Environment
            • Review flag
            • Vulnerability
            • Vulnerability Type
            • Attack Method
            • Risk Category
            • Framework
            • Assessment ID
            • Prompt Alias
            • Prompt Version
            • Prompt Label
            • Prompt Commit Hash
            • Prompt
            • Annotations
            • Status Code
            • Actor Type
          • conditionenum | enum | enum | enum | enum | enum | enum | enum | enum | enum

            Show 10 variantsHide 10 variants
            • enum

              Show 6 enum valuesHide 6 enum values
              • Is less than
              • Is equal or less than
              • Is greater than
              • Is equal or greater than
              • Is equal to
              • Does not equal
            • OR
            • enum

              Show 2 enum valuesHide 2 enum values
              • Has
              • Has not
            • OR
            • enum

              Show 2 enum valuesHide 2 enum values
              • Is
              • Is not
            • OR
            • enum

              Show 2 enum valuesHide 2 enum values
              • Is one of
              • Is not one of
            • OR
            • enum

              Show 4 enum valuesHide 4 enum values
              • Is
              • Is not
              • Is empty
              • Is not empty
            • OR
            • enum

              Show 2 enum valuesHide 2 enum values
              • Contains
              • Does not contain
            • OR
            • enum

              Show 3 enum valuesHide 3 enum values
              • Contains
              • Contains only
              • Does not contain
            • OR
            • enum

              Show 4 enum valuesHide 4 enum values
              • Has decreased by more than
              • Has decreased by less than
              • Has increased by more than
              • Has increased by less than
            • OR
            • enum

              Show 1 enum valueHide 1 enum value
              • Has changed from
            • OR
            • enum

              Show 1 enum valueHide 1 enum value
              • Is between
          • valuestring | number | list of strings

            Show 3 variantsHide 3 variants
            • string

            • OR
            • number

            • OR
            • list of strings

          • keystring

    • threadTimelimitinteger | null

      For THREAD rules, the seconds of inactivity waited before a thread is evaluated, or null when the rule uses the project's thread timelimit.

    • overwriteEvalsboolean

      Whether items that already have results for this metric collection are re-evaluated.

    • metricCollectionIdstring

      The id of the metric collection the rule runs. Retrieve it from the metric collections endpoint to see the metrics it holds.

    • createdAtstring

      When the rule was created, as an ISO 8601 datetime.

    • updatedAtstring

      When the rule was last changed, as an ISO 8601 datetime.

  • linkstring

    This is the URL of the resource on the Confident AI platform.

  • deprecatedboolean

    Indicates if this endpoint is deprecated.

Built byConfident AI