Log Prompts
Log prompts to LLM spans for version tracking in production
Overview
When you use prompts managed on Confident AI, you can log the exact prompt version used in each LLM call. Prompt logging works by:
- Pulling a prompt from Confident AI
- Recording which prompt version was used on the LLM span
That's it! This lets you monitor what prompts are running in production and which prompts perform best over time — because every trace (and every online eval result on it) is tied back to the prompt version that produced it, you can compare versions on real traffic instead of guessing.
Log a Prompt
Prompt logging is only meaningful for LLM spans. Make sure the span wrapping your model call has type="llm" set.
Pull and interpolate your prompt
Pull the prompt version from Confident AI and interpolate any variables.
main.py from deepeval.prompt import Prompt prompt = Prompt(alias="YOUR-PROMPT-ALIAS") prompt.pull() interpolated_prompt = prompt.interpolate(name="Joe")src/index.ts import { Prompt } from "deepeval"; const prompt = new Prompt({ alias: "YOUR-PROMPT-ALIAS" }); await prompt.pull(); const interpolatedPrompt = prompt.interpolate({ name: "Joe" });Use the prompt and record it on the span
Inside an LLM span, use the interpolated prompt for generation and record the prompt's alias and version on the span so you can see — and filter by — which prompt produced each call.
main.py from confident_trace import span, update_span from deepeval.prompt import Prompt from openai import OpenAI client = OpenAI() @span(type="llm", model="gpt-4o", provider="openai") def generate_response(user_input: str) -> str: prompt = Prompt(alias="YOUR-PROMPT-ALIAS") prompt.pull(version="00.00.01") interpolated_prompt = prompt.interpolate(name="Joe") response = client.chat.completions.create( model="gpt-4o", messages=interpolated_prompt, ) update_span( metadata={"prompt_alias": "YOUR-PROMPT-ALIAS", "prompt_version": "00.00.01"} ) return response.choices[0].message.contentsrc/index.ts import { span, updateSpan } from "confident-trace"; import { Prompt } from "deepeval"; import OpenAI from "openai"; const openai = new OpenAI(); const generateResponse = span( { name: "generate_response", type: "llm", model: "gpt-4o", provider: "openai" }, async (userInput: string) => { const prompt = new Prompt({ alias: "YOUR-PROMPT-ALIAS" }); await prompt.pull({ version: "00.00.01" }); const interpolatedPrompt = prompt.interpolate({ name: "Joe" }); const response = await openai.chat.completions.create({ model: "gpt-4o", messages: interpolatedPrompt as any[], }); updateSpan({ metadata: { promptAlias: "YOUR-PROMPT-ALIAS", promptVersion: "00.00.01" }, }); return response.choices[0].message.content; }, );
Once recorded, the prompt alias and version appear in the span's metadata in the trace view, making it easy to see exactly which prompt was used for each LLM call and to compare online eval results across prompt versions.
Next Steps
With prompts logged, set up cost tracking or refine what data your traces capture.
Track LLM Costs
Track token usage and cost for your LLM spans — manually or automatically.
Set Input/Output
Override the default input and output on traces and spans for better visualization and evaluation.
Last updated on