Launch Week 02 wrapped — explore all five launches

Update Metric

PUThttps://api.confident-ai.com/v2/metrics/{metricId}

Updates a custom metric and returns it. Only the fields you send are changed; send null to clear criteria or evaluationSteps, as long as one of them remains. Every update creates a new metric version.

PUT/v2/metrics/{metricId}
curl -X PUT "https://api.confident-ai.com/v2/metrics/{metricId}" \
  -H "CONFIDENT_API_KEY: <PROJECT-API-KEY>" \
  -H "Content-Type: application/json" \
  -d '{
  "criteria": "Determine if the actual output is correct based on the expected output.",
  "evaluationSteps": [
    "Compare the actual output with the expected output.",
    "Penalise any factual contradiction."
  ],
  "evaluationParams": [
    "actualOutput",
    "expectedOutput"
  ],
  "rubric": [
    {
      "scoreRange": [
        8,
        10
      ],
      "expectedOutcome": "The answer is factually correct and complete."
    }
  ]
}'
200
{
  "success": true,
  "data": {
    "id": "<METRIC-ID>",
    "name": "Correctness",
    "algorithm": "DEFAULT",
    "criteria": "Determine if the actual output is correct based on the expected output.",
    "evaluationSteps": null,
    "rubric": [
      {
        "scoreRange": [
          8,
          10
        ],
        "expectedOutcome": "The answer is factually correct and complete."
      }
    ],
    "dag": {
      "nodes": {
        "root": {
          "type": "BinaryJudgementNode",
          "criteria": "Does the actual output answer the input?",
          "evaluation_params": [
            "input",
            "actual_output"
          ],
          "children": [
            "pass",
            "fail"
          ]
        },
        "pass": {
          "type": "VerdictNode",
          "verdict": true,
          "score": 10
        },
        "fail": {
          "type": "VerdictNode",
          "verdict": false,
          "score": 0
        }
      }
    },
    "multiTurn": false,
    "requiredParameters": [
      "actualOutput",
      "expectedOutput"
    ]
  },
  "deprecated": false
}

Headers

  • CONFIDENT_API_KEYstringRequired

    The API key of your Confident AI project.

Path parameters

  • metricIdstringRequired

    The unique id of the metric.

Request body

  • criteriastring | null

    The new criteria, or null to clear it. One of criteria or evaluationSteps must remain set.

  • evaluationStepsarray | null

    The new evaluation steps, or null to clear them. One of criteria or evaluationSteps must remain set.

  • evaluationParamslist of enums

    The test case fields the metric evaluates. Each must match the metric's multiTurn.

    Show 14 enum valuesHide 14 enum values
    • input
    • actualOutput
    • expectedOutput
    • context
    • expectedTools
    • content
    • role
    • scenario
    • expectedOutcome
    • turns
    • toolsCalled
    • retrievalContext
    • metadata
    • tags
  • rubriclist of objects

    Score ranges that anchor how the metric scores.

    Show 2 propertiesHide 2 properties
    • scoreRangelist of anyRequired

      The inclusive start and end of the score range this outcome describes, each between 0 and 10, with the start no greater than the end.

    • expectedOutcomestringRequired

      What a response scoring in this range looks like.

Response

Update Metric succeeded.

  • successboolean

    Indicates if the request was successful.

  • dataobject

    A custom metric: how it scores, and which test case fields it needs.

    Show 9 propertiesHide 9 properties
    • idstring

      This is the unique id of the metric.

    • namestring

      This is the name of the metric, unique within your project.

    • algorithmenum | null

      The algorithm the metric is evaluated with. GEVAL scores against criteria or evaluation steps, DAG walks a decision graph, CODE runs your own code, and DEFAULT is a built-in metric.

      Show 4 enum valuesHide 4 enum values
      • DEFAULT
      • DAG
      • GEVAL
      • CODE
    • criteriastring | null

      This is the criteria the metric scores against, or null when it uses evaluation steps.

    • evaluationStepsarray | null

      These are the steps the metric follows to score, or null when it uses criteria.

    • rubricarray | null

      These are the score ranges that anchor how the metric scores, or null.

      Show 2 propertiesHide 2 properties
      • scoreRangelist of any

        The inclusive start and end of the score range this outcome describes, each between 0 and 10, with the start no greater than the end.

      • expectedOutcomestring

        What a response scoring in this range looks like.

    • dagobject | null

      The decision graph of a DAG metric. It is validated on write, and metric references are resolved against the metrics in your project.

      Show 1 propertyHide 1 property
      • nodesobject

        The graph's nodes keyed by node id, as serialized by deepeval's DAG metric. Judgement nodes list their children by id; verdict nodes carry a verdict and a score, or point at another metric by metric_name.

    • multiTurnboolean

      This is true when the metric evaluates conversations rather than single test cases.

    • requiredParameterslist of enums

      The test case fields the metric needs to run.

      Show 14 enum valuesHide 14 enum values
      • input
      • actualOutput
      • expectedOutput
      • context
      • expectedTools
      • content
      • role
      • scenario
      • expectedOutcome
      • turns
      • toolsCalled
      • retrievalContext
      • metadata
      • tags
  • deprecatedboolean

    Indicates if this endpoint is deprecated.

Built byConfident AI