Evaluation Task
Run a metric collection on incoming traces, spans, and threads, without code changes.
Overview
An evaluation task runs a metric collection on incoming traces, spans, or threads, without any code changes. It's how you set up online evaluations: pick what it runs on and which metrics, narrow it with filters and a sample rate, and every match is evaluated as it arrives. On the Workflows page, evaluation tasks are listed under Evaluation Rules.
Create an Evaluation Task
- Navigate to Workflows and select Traces, Spans, or Threads
- Expand Evaluation Rules in the panel beside the graph
- Click New rule
- Configure the task in the side drawer (see fields below)
- Click Create Rule
Fields
| Field | Required | Description |
|---|---|---|
| Name | Yes | A unique name for the task |
| Description | No | Optional context about the task's purpose |
| Enabled | Yes | Toggle on to activate; disabled tasks are saved but skipped at ingest time |
| Data Model | Yes | Trace, Span, or Thread — determines what the task runs on and when |
| Span Type | Span tasks only | Restrict to a specific span type: LLM, Agent, Tool, Retriever, or Custom. Leave as Any to match all spans. |
| Metric Collection | Yes | The metric collection to run. Trace and span tasks require a single-turn collection; thread tasks require a multi-turn collection. |
| Filters | No | Scope the task to a subset of data (e.g. specific environments, tags, or metadata values). Leave empty to match every entity. |
| Sample Rate | No | Fraction of matching entities the task fires on (0.0–1.0). Sampling is deterministic — the same item always makes the same decision for a given task. Defaults to 1.0. See Sample Rate for how collection and per-metric rates compound. |
| Time Limit | Thread tasks only | Seconds of inactivity before a thread is eligible for evaluation. The thread evaluates once no new traces have arrived for this period. Defaults to 300. |
| Overwrite Evaluations | Thread tasks only | When on, each idle cycle replaces the thread's prior evaluations. When off (default), each cycle appends a new set of metric rows, preserving the full history. |
Data models
| Data Model | When it runs | Metric collection type |
|---|---|---|
| Trace | At ingest, on each incoming trace | Single-turn |
| Span | At ingest, on each incoming span | Single-turn |
| Thread | After the thread has been idle for the configured time limit | Multi-turn |
Filters
Filters narrow which traces, spans, or threads a task applies to. Filters can target environment, tags, metadata fields, latency, and other dimensions. Filter tabs for eval metrics, annotations, and signals are not available here — those dimensions don't exist at ingest time.
Threads
For threads, evaluation tasks are the primary way to run evaluations automatically — there is no inline SDK parameter that triggers a thread-level evaluation. Threads can still be evaluated explicitly from code if needed; see Evaluate Threads.
Next Steps
Evaluate Traces & Spans
Log the test case parameters your metrics need, and see what each task evaluates.
Evaluate Threads
Evaluate whole conversations once they go idle.
Last updated on