Launch Week 3: Five days of launches

Evaluation Task

Run a metric collection on incoming traces, spans, and threads, without code changes.

Overview

An evaluation task runs a metric collection on incoming traces, spans, or threads, without any code changes. It's how you set up online evaluations: pick what it runs on and which metrics, narrow it with filters and a sample rate, and every match is evaluated as it arrives. On the Workflows page, evaluation tasks are listed under Evaluation Rules.

An evaluation task

Create an Evaluation Task

  1. Navigate to Workflows and select Traces, Spans, or Threads
  2. Expand Evaluation Rules in the panel beside the graph
  3. Click New rule
  4. Configure the task in the side drawer (see fields below)
  5. Click Create Rule

Fields

FieldRequiredDescription
NameYesA unique name for the task
DescriptionNoOptional context about the task's purpose
EnabledYesToggle on to activate; disabled tasks are saved but skipped at ingest time
Data ModelYesTrace, Span, or Thread — determines what the task runs on and when
Span TypeSpan tasks onlyRestrict to a specific span type: LLM, Agent, Tool, Retriever, or Custom. Leave as Any to match all spans.
Metric CollectionYesThe metric collection to run. Trace and span tasks require a single-turn collection; thread tasks require a multi-turn collection.
FiltersNoScope the task to a subset of data (e.g. specific environments, tags, or metadata values). Leave empty to match every entity.
Sample RateNoFraction of matching entities the task fires on (0.0–1.0). Sampling is deterministic — the same item always makes the same decision for a given task. Defaults to 1.0. See Sample Rate for how collection and per-metric rates compound.
Time LimitThread tasks onlySeconds of inactivity before a thread is eligible for evaluation. The thread evaluates once no new traces have arrived for this period. Defaults to 300.
Overwrite EvaluationsThread tasks onlyWhen on, each idle cycle replaces the thread's prior evaluations. When off (default), each cycle appends a new set of metric rows, preserving the full history.

Data models

Data ModelWhen it runsMetric collection type
TraceAt ingest, on each incoming traceSingle-turn
SpanAt ingest, on each incoming spanSingle-turn
ThreadAfter the thread has been idle for the configured time limitMulti-turn

Filters

Filters narrow which traces, spans, or threads a task applies to. Filters can target environment, tags, metadata fields, latency, and other dimensions. Filter tabs for eval metrics, annotations, and signals are not available here — those dimensions don't exist at ingest time.

Threads

For threads, evaluation tasks are the primary way to run evaluations automatically — there is no inline SDK parameter that triggers a thread-level evaluation. Threads can still be evaluated explicitly from code if needed; see Evaluate Threads.

Next Steps

Ready to monitor AI in production?Connect traces, alerts, dashboards, and evals in one production workflowBook a demo

Last updated on

Built byConfident AI