Launch Week 3: Five days of launches

Prompt

Write your own judge prompt and fill it with data from your traces and spans

Overview

A prompt metric is a custom LLM-as-a-judge metric where you write the judge prompt yourself. You fill the prompt with data from the trace or span being evaluated using variables, and the judge returns a score and a reason.

Use it when you already have a judge prompt that works, such as one you've been running in another tool, and want Confident AI to run it exactly as written.

Why Prompt?

G-Eval turns your criteria into evaluation steps and wraps them in its own judge prompt. A prompt metric skips that step:

  • Your prompt, as written: the judge sees the messages you wrote with the variables filled in, and nothing is rewritten or generated from them
  • Reads any part of your data: a variable can point at any value inside a trace or span's input, output, metadata, or tool calls, such as the last message in an OpenAI messages array
  • Checked before you save: you can preview what every variable resolves to, and run the judge on a real trace or span from your project

Use a prompt metric when you know exactly what the judge should be told. Use G-Eval when you would rather describe a criteria in plain language and let Confident AI write the evaluation steps.

How It Works

A prompt metric works in a few simple steps:

  1. Fills each {{variable}} in your messages with the value its path points to in the trace or span being evaluated
  2. Sends the messages to your evaluation model, along with an instruction to return a score from 0 - 10 and a reason
  3. Divides the score by 10 to normalize it to the range 0 - 1

Writing the Prompt

The prompt is a list of messages, and each message has a role:

  • System: instructions for the judge, such as who it is and how strict to be. A system message can only be the first message.
  • User: the content to judge, usually where most of your variables go.
  • Assistant: an example of how the judge should answer, if you want to show it one.

Write {{name}} anywhere in a message to insert a value, where name is made of letters, numbers, and underscores. For example, a prompt that judges a payroll help assistant could look like this:

QUESTION:
{{question}}

ANSWER:
{{answer}}

TOOLS USED:
{{tools}}

Give 10 if the answer fully solves the question. If the answer declines to help, it must say why and point to a next step. Give 0 if the answer makes up facts that the tools used do not support.

Mapping Variables

Every variable in your prompt needs a path. A path starts with a field, then walks into that field one step at a time.

FieldWhat it reads
InputThe input of the trace or span
OutputThe output of the trace or span
MetadataThe metadata of the trace or span
Tools CalledThe tools called, as a list where each tool call has a name, input_parameters, and output

After the field, each step is one of:

  • A key, such as messages or content, which reads a value from an object
  • A position, such as 0 for the first item or -1 for the last item, which reads one item from a list
  • all, which reads every item in a list

For the payroll prompt above, the paths could be:

VariablePathWhat it gets
{{question}}Input › messages › -1 › contentThe user's last message
{{answer}}Output › messages › -1 › contentThe assistant's reply
{{tools}}Tools Called › all › nameThe names of every tool that was called

A path that stops at the field, such as just Input, inserts the whole field. This is what you want when your input and output are plain text rather than OpenAI-style messages.

Create a Prompt Metric via the UI

Prompt metrics can be created under Project > Metrics > Library.

  1. Fill in metric details

    Provide the metric name, and optionally a description, then select Prompt as the algorithm. Prompt metrics are single-turn only, so the multi-turn toggle is turned off.

  2. Write the judge prompt

    Write your first message, and click Add Message to add more. Each {{variable}} you type shows up in the Variables section below, ready to be mapped.

  3. Map variables

    Each variable has a path bar. Pick the field first, then type a key or a position, or pick one from the suggestions. You can also paste a full path, such as messages[-1].content, and it will be split into steps for you.

    To change a step, click it or move to it with the arrow keys, then type the new value in place.

  4. Test on a sample (optional)

    Pick a trace or span from your project under Test on Sample. Every variable then shows a preview of its value, and the path bar suggests the keys and positions that exist in your sample.

    Click Run test to run the judge on that sample and see the score and reason it returns. Nothing is saved, but the test uses your evaluation model.

  5. Review and save

    Make sure your prompt and variables look right in the final review page, and click Save.

Using Prompt Metrics

Once saved, a prompt metric behaves like any other metric on Confident AI. Add it to a metric collection, where you can set its threshold and strictness, to use it in evaluations.

Prompt metrics always run on your evaluation model, so their eval mode is fixed to LLM and cannot be changed in a metric collection.

Scaling beyond prototype?For teams evaluating Confident AI in productionTalk to us

Last updated on

Built byConfident AI