Launch Week 3: Five days of launches

Dataset Ingestion Tasks

Overview

The Confident AI SDK exposes every Dataset Ingestion Task method on the platform. This page documents how to call these methods in all supported languages. See the introduction to install the SDK and set your API key.

Methods

List Ingestion Tasks

Lists the ingestion tasks on the dataset as summary rows. Retrieve one by id for its full configuration. Requires the Starter plan or above.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.list_ingestion_tasks(
    dataset_id="<DATASET-ID>",
    data_model="TRACE",
)

For async mode, call a_list_ingestion_tasks and await it as shown below:

result = await client.datasets.a_list_ingestion_tasks(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
data_modelOptional[Literal['TRACE', 'SPAN', 'THREAD']]Only return tasks harvesting this kind of item. Omit it to return all of them.

Returns

This method returns an object of type DatasetIngestionTaskList.

Create Ingestion Task

Creates a standing rule that harvests matching production traces, spans or threads into the dataset as goldens, starting immediately unless enabled is false, and returns its id. dataModel must match the dataset: THREAD for multi-turn, TRACE or SPAN for single-turn. Requires the Starter plan or above.

from confident_ai import ConfidentAI
from confident_ai.common import IngestionDataModel

client = ConfidentAI()

result = client.datasets.create_ingestion_task(
    dataset_id="<DATASET-ID>",
    name="Harvest failed lookups",
    data_model=IngestionDataModel.TRACE,
    description="Traces where the assistant failed to name a capital.",
    enabled=True,
    sample_rate=0.1,
    filters={
        "operator": "AND",
        "groups": [
            {
                "operator": "AND",
                "filters": [
                    {
                        "category": "Name",
                        "condition": "Is",
                        "value": "capital-lookup"
                    }
                ]
            }
        ]
    },
    max_goldens=500,
    input_transformer_id="<TRANSFORMER-ID>",
    output_transformer_id="<OUTPUT-TRANSFORMER-ID>",
    include_input=True,
    include_actual_output=True,
    include_expected_output=False,
    include_retrieval_context=False,
    include_context=False,
    include_tools_called=False,
    include_expected_tools=False,
)

For async mode, call a_create_ingestion_task and await it as shown below:

result = await client.datasets.a_create_ingestion_task(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
namestrRequired. A name for the task, unique within the dataset.
data_modelIngestionDataModelRequired. See IngestionDataModel.
descriptionOptional[str]A note about what the task harvests. Send null to clear it.
enabledOptional[bool]Whether the task runs. Disabling it unschedules the harvesting job, and goldens already created are kept. Defaults to false.
sample_rateOptional[float]The fraction of matching items to ingest, between 0 and 1. Defaults to 1, all of them.
filtersOptional[FilterSet]See FilterSet.
max_goldensOptional[int]The maximum number of goldens this task will ever create. Send null to remove the cap.
input_transformer_idOptional[str]The id of a transformer that reshapes the harvested input before it is stored. Send null to detach it.
output_transformer_idOptional[str]The id of a transformer that reshapes the harvested output before it is stored. Send null to detach it.
include_inputOptional[bool]Populate the golden's input from the harvested item. Defaults to true; every other include flag defaults to false.
include_actual_outputOptional[bool]Populate the golden's actualOutput from the harvested item.
include_expected_outputOptional[bool]Populate the golden's expectedOutput from the harvested item.
include_retrieval_contextOptional[bool]Populate the golden's retrievalContext from the harvested item.
include_contextOptional[bool]Populate the golden's context from the harvested item.
include_tools_calledOptional[bool]Populate the golden's toolsCalled from the harvested item.
include_expected_toolsOptional[bool]Populate the golden's expectedTools from the harvested item.

Returns

This method returns an object of type DatasetIngestionTaskRef.

Get Ingestion Task

Retrieves a single ingestion task on the dataset, with its full configuration.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.get_ingestion_task(
    dataset_id="<DATASET-ID>",
    dataset_ingestion_task_id="<DATASET-INGESTION-TASK-ID>",
)

For async mode, call a_get_ingestion_task and await it as shown below:

result = await client.datasets.a_get_ingestion_task(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
dataset_ingestion_task_idstrRequired. The unique id of the ingestion task.

Returns

This method returns an object of type DatasetIngestionTask.

Update Ingestion Task

Updates an ingestion task and returns it. Only the fields you send are changed, and at least one is required; send null to clear a nullable field. Toggling enabled schedules or unschedules the harvesting job.

from confident_ai import ConfidentAI
from confident_ai.common import IngestionDataModel

client = ConfidentAI()

result = client.datasets.update_ingestion_task(
    dataset_id="<DATASET-ID>",
    dataset_ingestion_task_id="<DATASET-INGESTION-TASK-ID>",
    name="Harvest failed lookups",
    data_model=IngestionDataModel.TRACE,
    description="Traces where the assistant failed to name a capital.",
    enabled=True,
    sample_rate=0.1,
    filters={
        "operator": "AND",
        "groups": [
            {
                "operator": "AND",
                "filters": [
                    {
                        "category": "Name",
                        "condition": "Is",
                        "value": "capital-lookup"
                    }
                ]
            }
        ]
    },
    max_goldens=500,
    input_transformer_id="<TRANSFORMER-ID>",
    output_transformer_id="<OUTPUT-TRANSFORMER-ID>",
    include_input=True,
    include_actual_output=True,
    include_expected_output=False,
    include_retrieval_context=False,
    include_context=False,
    include_tools_called=False,
    include_expected_tools=False,
)

For async mode, call a_update_ingestion_task and await it as shown below:

result = await client.datasets.a_update_ingestion_task(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
dataset_ingestion_task_idstrRequired. The unique id of the ingestion task.
nameOptional[str]A new name for the task, unique within the dataset.
data_modelOptional[IngestionDataModel]See IngestionDataModel.
descriptionOptional[str]A note about what the task harvests. Send null to clear it.
enabledOptional[bool]Whether the task runs. Disabling it unschedules the harvesting job, and goldens already created are kept. Defaults to false.
sample_rateOptional[float]The fraction of matching items to ingest, between 0 and 1. Defaults to 1, all of them.
filtersOptional[FilterSet]See FilterSet.
max_goldensOptional[int]The maximum number of goldens this task will ever create. Send null to remove the cap.
input_transformer_idOptional[str]The id of a transformer that reshapes the harvested input before it is stored. Send null to detach it.
output_transformer_idOptional[str]The id of a transformer that reshapes the harvested output before it is stored. Send null to detach it.
include_inputOptional[bool]Populate the golden's input from the harvested item. Defaults to true; every other include flag defaults to false.
include_actual_outputOptional[bool]Populate the golden's actualOutput from the harvested item.
include_expected_outputOptional[bool]Populate the golden's expectedOutput from the harvested item.
include_retrieval_contextOptional[bool]Populate the golden's retrievalContext from the harvested item.
include_contextOptional[bool]Populate the golden's context from the harvested item.
include_tools_calledOptional[bool]Populate the golden's toolsCalled from the harvested item.
include_expected_toolsOptional[bool]Populate the golden's expectedTools from the harvested item.

Returns

This method returns an object of type DatasetIngestionTask.

Delete Ingestion Task

Permanently deletes an ingestion task and unschedules its harvesting job. Goldens it already created stay in the dataset. This action cannot be undone.

from confident_ai import ConfidentAI

client = ConfidentAI()

result = client.datasets.delete_ingestion_task(
    dataset_id="<DATASET-ID>",
    dataset_ingestion_task_id="<DATASET-INGESTION-TASK-ID>",
)

For async mode, call a_delete_ingestion_task and await it as shown below:

result = await client.datasets.a_delete_ingestion_task(...)

Parameters

ParameterTypeDescription
dataset_idstrRequired. The unique id of the dataset.
dataset_ingestion_task_idstrRequired. The unique id of the ingestion task.

Returns

This method returns an object of type DatasetIngestionTaskRef.

Types

DatasetIngestionTask

An ingestion task: which production items it harvests into the dataset, how it samples them, and which golden fields it fills.

class DatasetIngestionTask:
    id: str
    name: str
    description: Optional[str]
    enabled: bool
    sample_rate: float = Field(alias="sampleRate")
    data_model: IngestionDataModel = Field(alias="dataModel")
    filters: Optional[FilterSet]
    max_goldens: Optional[int] = Field(alias="maxGoldens")
    input_transformer_id: Optional[str] = Field(alias="inputTransformerId")
    output_transformer_id: Optional[str] = Field(alias="outputTransformerId")
    include_input: bool = Field(alias="includeInput")
    include_actual_output: bool = Field(alias="includeActualOutput")
    include_expected_output: bool = Field(alias="includeExpectedOutput")
    include_retrieval_context: bool = Field(alias="includeRetrievalContext")
    include_context: bool = Field(alias="includeContext")
    include_tools_called: bool = Field(alias="includeToolsCalled")
    include_expected_tools: bool = Field(alias="includeExpectedTools")
    created_at: str = Field(alias="createdAt")
    updated_at: str = Field(alias="updatedAt")

idstrRequired

The unique id of the ingestion task.

Example: "<DATASET-INGESTION-TASK-ID>"

namestrRequired

The name of the ingestion task, unique within the dataset.

Example: "Harvest failed lookups"

descriptionOptional[str]Required

A note about what the task harvests, or null.

Example: "Traces where the assistant failed to name a capital."

enabledboolRequired

Whether the task is currently harvesting.

Example: true

sample_ratefloatRequired

The fraction of matching items the task ingests, between 0 and 1.

Example: 0.1

data_modelIngestionDataModelRequired

filtersOptional[FilterSet]Required

max_goldensOptional[int]Required

The maximum number of goldens this task will ever create, or null when uncapped.

Example: 500

input_transformer_idOptional[str]Required

The id of the transformer that reshapes the harvested input, or null when the task does not use one.

Example: "<TRANSFORMER-ID>"

output_transformer_idOptional[str]Required

The id of the transformer that reshapes the harvested output, or null when the task does not use one.

include_inputboolRequired

Whether the golden's input is populated from the harvested item.

Example: true

include_actual_outputboolRequired

Whether the golden's actualOutput is populated from the harvested item.

Example: true

include_expected_outputboolRequired

Whether the golden's expectedOutput is populated from the harvested item.

Example: false

include_retrieval_contextboolRequired

Whether the golden's retrievalContext is populated from the harvested item.

Example: false

include_contextboolRequired

Whether the golden's context is populated from the harvested item.

Example: false

include_tools_calledboolRequired

Whether the golden's toolsCalled is populated from the harvested item.

Example: false

include_expected_toolsboolRequired

Whether the golden's expectedTools is populated from the harvested item.

Example: false

created_atstrRequired

When the task was created, as an ISO 8601 timestamp.

Example: "2026-05-28T13:05:24.777000+00:00"

updated_atstrRequired

When the task was last updated, as an ISO 8601 timestamp.

Example: "2026-05-28T13:35:16.268000+00:00"

DatasetIngestionTaskList

class DatasetIngestionTaskList:
    dataset_ingestion_tasks: List[DatasetIngestionTaskSummary] = Field(alias="datasetIngestionTasks")

dataset_ingestion_tasksList[DatasetIngestionTaskSummary]Required

This is the list of ingestion tasks on the dataset, newest first, as summary rows.

See DatasetIngestionTaskSummary.

DatasetIngestionTaskRef

class DatasetIngestionTaskRef:
    id: str

idstrRequired

The unique id of the ingestion task.

Example: "<DATASET-INGESTION-TASK-ID>"

DatasetIngestionTaskSummary

A row in the dataset's ingestion task list. Its full configuration comes from the get endpoint.

class DatasetIngestionTaskSummary:
    id: str
    name: str
    enabled: bool
    data_model: IngestionDataModel = Field(alias="dataModel")

idstrRequired

The unique id of the ingestion task.

Example: "<DATASET-INGESTION-TASK-ID>"

namestrRequired

The name of the ingestion task, unique within the dataset.

Example: "Harvest failed lookups"

enabledboolRequired

Whether the task is currently harvesting.

Example: true

data_modelIngestionDataModelRequired

FilterSet

A set of filter groups combined by a top-level operator. Each group combines its filter rows by its own operator, and each row matches one property, such as Name or User Id, against a value with a condition such as Is or Contains.

class FilterSet:
    operator: Literal["AND", "OR"]
    groups: List[FilterSetGroup]

operatorLiteral["AND", "OR"]Required

groupsList[FilterSetGroup]Required

IngestionDataModel

What kind of production item an ingestion task harvests. THREAD tasks fill multi-turn datasets; TRACE and SPAN tasks fill single-turn ones.

class IngestionDataModel(Enum):
    TRACE = "TRACE"
    SPAN = "SPAN"
    THREAD = "THREAD"

TRACE · SPAN · THREAD

Building a production pipeline?Design a scalable API workflow for evals, datasets, traces, and promptsTalk to an engineer

Last updated on

Built byConfident AI