Launch Week 3: Five days of launches

Track LLM Costs

Track the token usage and cost of your LLM calls

Overview

Confident AI tracks the token usage and cost of your LLM calls, helping you identify high-cost models and heavy usage patterns across your application.

LLM Cost Tracking

How It Works

Confident AI resolves token usage and cost for each LLM span in the following order of precedence, and separately for both input and output tokens:

  1. Per-token costs and counts set in confident-trace (via span(...) or update_span() / updateSpan()) take the highest priority and will always override any other source.
    • Integrations may provide the token count, but cost calculation happens the same way.
  2. Custom set model costs — if you provide token counts but not per-token costs, Confident AI will use the pricing you've configured in your Model Costs settings to calculate the cost.
  3. Automatic inference — if neither per-token costs nor project-level costs are available, Confident AI tokenizes the span's input/output text using a provider-specific tokenizer and internally looks up pricing based on the model.

Automatic Usage Capture

If you're using the OpenAI integration or any other supported provider integration, you don't need to do anything: init() instruments the client and each call's input_token_count and output_token_count are captured on the LLM span using the standard OpenTelemetry GenAI attributes.

Track Token Usage Count

You can manually set the input and output token counts on an LLM span using update_span() / updateSpan(). This is useful when your provider returns token usage in the response and you want to log it precisely, or when you're calling a model no integration covers.

main.py
from confident_trace import span, update_span

@span(type="llm", model="gpt-4o", provider="openai")
def generate_response(prompt: str) -> str:
    response = call_llm(prompt)
    update_span(
        input_token_count=response.usage.prompt_tokens,
        output_token_count=response.usage.completion_tokens,
    )
    return response.text

Token counts must be non-negative integers, and an explicit 0 is kept as 0 rather than treated as unset.

If you don't provide token counts and aren't using an integration, Confident AI will attempt to infer them by tokenizing the span's input and output text using the appropriate provider tokenizer. The table below summarizes each supported provider and its tokenization method.

ProviderTokenizerExample ModelsToken Counting Method
OpenAItiktokengpt-4o, gpt-4.1, o1, o3Client-side tokenization using model-specific encodings
Anthropic@anthropic-ai/tokenizerclaude-3.5-sonnet, claude-3.7-sonnet, claude-4Claude-specific tokenization algorithm
GoogleGemini APIgemini-2.0-flash, gemini-2.5-proServer-side token counting via API call

See the OpenAI documentation, Anthropic documentation, or Google documentation for the most up-to-date pricing.

Track Token Usage Cost

Once token counts are available (either set manually, captured by an integration, or inferred automatically), Confident AI resolves the per-token cost using the following precedence:

  1. Per-token costs set in code — if you provide cost per input/output tokens directly via span(...) or update_span() / updateSpan(), these always take priority.
  2. Custom set model costs — if per-token costs aren't set in code, Confident AI uses the pricing configured in your Model Costs settings.
  3. Automatic price lookup — if no project-level costs are configured, Confident AI looks up the per-token pricing internally based on the model. This is only available for OpenAI, Anthropic, and Gemini models.

If none of the above resolve a per-token cost, the cost for that side (input or output) is not logged.

Explicit Cost Setting

Set the per-token costs explicitly on the span alongside your token counts. This spares you for provider models not supported by automatic price lookup — self-hosted models, custom gateways, or negotiated enterprise pricing.

main.py
from confident_trace import span, update_span

@span(
    type="llm", model="my-model", provider="custom",
    cost_per_input_token=0.000001,
    cost_per_output_token=0.000002,
)
def generate_response(prompt: str) -> str:
    response = call_llm(prompt)
    update_span(
        input_token_count=response.usage.prompt_tokens,
        output_token_count=response.usage.completion_tokens,
    )
    return response.text

You can pass the cost fields either on the span options (as above) or to update_span() / updateSpan() — whichever is more convenient. Either way they only apply to an LLM span; on any other span type the general fields are still applied but the LLM fields are skipped with a one-time warning.

Custom Price Lookup

If you provide token counts but don't set per-token costs in code, Confident AI will use the pricing you've configured in your project's Model Costs settings. This is useful when you want to manage pricing centrally without changing any code.

Model costs are matched against the model name on your LLM span using case-insensitive regex patterns, which match anywhere in the name unless anchored. For example:

  • ^gpt-4o$ — matches only gpt-4o
  • gpt-4.* — matches gpt-4o, gpt-4o-mini, gpt-4-turbo, etc.
  • claude-.* — matches all Claude model variants

You can optionally restrict a cost rule to a specific provider, and set input and output costs independently per million tokens. See the full Model Costs settings page for setup instructions.

Configure Model Costs

Automatic Price Lookup

If you provide a supported model on your LLM span and neither SDK-level nor project-level costs are configured, Confident AI will automatically look up the per-token pricing and calculate the cost — no additional code needed.

main.py
from confident_trace import span, update_span

@span(type="llm", model="gpt-4o", provider="openai")
def generate_response(prompt: str) -> str:
    output = call_llm(prompt)
    update_span(input=prompt, output=output)
    return output

Cost on Traces

Cost on traces are automatically set by summing up the cost of all LLM spans in said trace. Similar to LLM spans, trace cost defaults to null values if no LLM spans have non-null values.

Next Steps

With cost tracking configured, continue setting up the rest of your instrumentation.

Ready to monitor AI in production?Connect traces, alerts, dashboards, and evals in one production workflowBook a demo

Last updated on

Built byConfident AI