Launch Week 3: Five days of launches

Link Test Cases to Traces

Link evaluation test cases and turns to their traces for full observability.

Overview

When you run evaluations through an AI Connection, Confident AI can link each result back to the trace your AI app produced. This gives you full observability—jump straight from an evaluation result to the exact trace that generated it.

There are two flavors of trace linking:

  • Linking test cases to traces for single-turn evaluations
  • Linking turns to traces for multi-turn evaluations and red-team attacks

Both work by passing an identifier (testCaseId or turnId) from your payload into your tracing setup. With confident-trace, use trace_context / traceContext to supply the identifier before an instrumented call starts, or update_trace / updateTrace if a custom span has already started the trace.

Linking Test Cases to Traces

For single-turn evaluations, you can link each test case to its corresponding trace for full observability. This is done by including testCaseId in your payload (enabled by default) and passing it to your tracing setup.

sequenceDiagram
    participant C as Confident AI
    participant E as Your Endpoint
    participant T as Tracing

    C->>E: Ping AI Connection with testCaseId
    E->>T: Create trace with testCaseId
    E-->>C: Return actual_output
    T-->>C: Send trace to Confident AI
    Note over C: Trace linked to test case

Include testCaseId in your payload configuration and ensure your AI connection is configured to accept it.

{
  "input": golden.input,
  "testCaseId": testCaseId
}

Because testCaseId is available before your app starts its traced work, pass it through a trace_context / traceContext. The context does not create a trace or an extra span. It supplies the ID to the trace created by the auto-instrumented integration call inside it.

Each example uses a FastAPI request handler and calls init() once when the server starts.

from fastapi import FastAPI
from pydantic import BaseModel
from langchain_openai import ChatOpenAI
from confident_trace import init, trace_context

init()
app = FastAPI()
model = ChatOpenAI(model="gpt-4o")

class GenerateRequest(BaseModel):
    input: str
    testCaseId: str

@app.post("/generate")
def generate(request: GenerateRequest):
    with trace_context(test_case_id=request.testCaseId):
        output = model.invoke(request.input).content
    return {"output": output}

Linking Turns to Traces

For multi-turn evaluations and multi-turn red-team attacks, Confident AI calls your endpoint once per turn. Each turn has its own turnId that you can pass to your tracing setup. This links each turn's trace to the specific turn in the conversation, letting you view traces per-turn from the evaluation or assessment results.

sequenceDiagram
    participant C as Confident AI
    participant E as Your Endpoint
    participant T as Tracing

    Note over C,E: Turn 1
    C->>E: { turnId, ...payload }
    E->>T: Create trace with turnId
    E-->>C: Return actual_output

    Note over C,E: Turn 2
    C->>E: { turnId, ...payload }
    E->>T: Create trace with turnId
    E-->>C: Return actual_output

    T-->>C: Send traces to Confident AI
    Note over C: Each turn linked to its trace

Include turnId (alongside testCaseId) in your payload configuration:

{
  "input": golden.input,
  "testCaseId": testCaseId,
  "turnId": turnId,
  "state": state
}

Then, supply both IDs to the trace created for each request:

Each example uses a FastAPI request handler and calls init() once when the server starts.

from fastapi import FastAPI
from pydantic import BaseModel
from langchain_openai import ChatOpenAI
from confident_trace import init, trace_context

init()
app = FastAPI()
model = ChatOpenAI(model="gpt-4o")

class GenerateRequest(BaseModel):
    input: str
    testCaseId: str
    turnId: str

@app.post("/generate")
def generate(request: GenerateRequest):
    with trace_context(
        test_case_id=request.testCaseId,
        turn_id=request.turnId,
    ):
        output = model.invoke(request.input).content
    return {"output": output}

Next Steps

With traces linked to your evaluation results, you can debug failures end-to-end. Explore related observability and connection features next.

Scaling beyond prototype?For teams evaluating Confident AI in productionTalk to us

Last updated on

Built byConfident AI