Prompt
Write your own judge prompt and fill it with data from your traces and spans
Overview
A prompt metric is a custom LLM-as-a-judge metric where you write the judge prompt yourself. You fill the prompt with data from the trace or span being evaluated using variables, and the judge returns a score and a reason.
Use it when you already have a judge prompt that works, such as one you've been running in another tool, and want Confident AI to run it exactly as written.
Why Prompt?
G-Eval turns your criteria into evaluation steps and wraps them in its own judge prompt. A prompt metric skips that step:
- Your prompt, as written: the judge sees the messages you wrote with the variables filled in, and nothing is rewritten or generated from them
- Reads any part of your data: a variable can point at any value inside a trace or span's input, output, metadata, or tool calls, such as the last message in an OpenAI
messagesarray - Checked before you save: you can preview what every variable resolves to, and run the judge on a real trace or span from your project
Use a prompt metric when you know exactly what the judge should be told. Use G-Eval when you would rather describe a criteria in plain language and let Confident AI write the evaluation steps.
How It Works
A prompt metric works in a few simple steps:
- Fills each
{{variable}}in your messages with the value its path points to in the trace or span being evaluated - Sends the messages to your evaluation model, along with an instruction to return a score from 0 - 10 and a reason
- Divides the score by 10 to normalize it to the range 0 - 1
Writing the Prompt
The prompt is a list of messages, and each message has a role:
- System: instructions for the judge, such as who it is and how strict to be. A system message can only be the first message.
- User: the content to judge, usually where most of your variables go.
- Assistant: an example of how the judge should answer, if you want to show it one.
Write {{name}} anywhere in a message to insert a value, where name is made of letters, numbers, and underscores. For example, a prompt that judges a payroll help assistant could look like this:
QUESTION:
{{question}}
ANSWER:
{{answer}}
TOOLS USED:
{{tools}}
Give 10 if the answer fully solves the question. If the answer declines to help, it must say why and point to a next step. Give 0 if the answer makes up facts that the tools used do not support.Mapping Variables
Every variable in your prompt needs a path. A path starts with a field, then walks into that field one step at a time.
| Field | What it reads |
|---|---|
| Input | The input of the trace or span |
| Output | The output of the trace or span |
| Metadata | The metadata of the trace or span |
| Tools Called | The tools called, as a list where each tool call has a name, input_parameters, and output |
After the field, each step is one of:
- A key, such as
messagesorcontent, which reads a value from an object - A position, such as
0for the first item or-1for the last item, which reads one item from a list all, which reads every item in a list
For the payroll prompt above, the paths could be:
| Variable | Path | What it gets |
|---|---|---|
{{question}} | Input › messages › -1 › content | The user's last message |
{{answer}} | Output › messages › -1 › content | The assistant's reply |
{{tools}} | Tools Called › all › name | The names of every tool that was called |
A path that stops at the field, such as just Input, inserts the whole field. This is what you want when your input and output are plain text rather than OpenAI-style messages.
Create a Prompt Metric via the UI
Prompt metrics can be created under Project > Metrics > Library.
Fill in metric details
Provide the metric name, and optionally a description, then select Prompt as the algorithm. Prompt metrics are single-turn only, so the multi-turn toggle is turned off.
Write the judge prompt
Write your first message, and click Add Message to add more. Each
{{variable}}you type shows up in the Variables section below, ready to be mapped.Map variables
Each variable has a path bar. Pick the field first, then type a key or a position, or pick one from the suggestions. You can also paste a full path, such as
messages[-1].content, and it will be split into steps for you.To change a step, click it or move to it with the arrow keys, then type the new value in place.
Test on a sample (optional)
Pick a trace or span from your project under Test on Sample. Every variable then shows a preview of its value, and the path bar suggests the keys and positions that exist in your sample.
Click Run test to run the judge on that sample and see the score and reason it returns. Nothing is saved, but the test uses your evaluation model.
Review and save
Make sure your prompt and variables look right in the final review page, and click Save.
Using Prompt Metrics
Once saved, a prompt metric behaves like any other metric on Confident AI. Add it to a metric collection, where you can set its threshold and strictness, to use it in evaluations.
Prompt metrics always run on your evaluation model, so their eval mode is fixed to LLM and cannot be changed in a metric collection.
Scaling beyond prototype?For teams evaluating Confident AI in productionTalk to usLast updated on