Track LLM Costs
Track the token usage and cost of your LLM calls
Overview
Confident AI tracks the token usage and cost of your LLM calls, helping you identify high-cost models and heavy usage patterns across your application.
How It Works
Confident AI resolves token usage and cost for each LLM span in the following order of precedence, and separately for both input and output tokens:
- Per-token costs and counts set in
confident-trace(viaspan(...)orupdate_span()/updateSpan()) take the highest priority and will always override any other source.- Integrations may provide the token count, but cost calculation happens the same way.
- Custom set model costs — if you provide token counts but not per-token costs, Confident AI will use the pricing you've configured in your Model Costs settings to calculate the cost.
- Automatic inference — if neither per-token costs nor project-level costs are available, Confident AI tokenizes the span's input/output text using a provider-specific tokenizer and internally looks up pricing based on the
model.
Automatic Usage Capture
If you're using the OpenAI integration or any other supported provider integration, you don't need to do anything: init() instruments the client and each call's input_token_count and output_token_count are captured on the LLM span using the standard OpenTelemetry GenAI attributes.
Track Token Usage Count
You can manually set the input and output token counts on an LLM span using update_span() / updateSpan(). This is useful when your provider returns token usage in the response and you want to log it precisely, or when you're calling a model no integration covers.
from confident_trace import span, update_span
@span(type="llm", model="gpt-4o", provider="openai")
def generate_response(prompt: str) -> str:
response = call_llm(prompt)
update_span(
input_token_count=response.usage.prompt_tokens,
output_token_count=response.usage.completion_tokens,
)
return response.textimport { span, updateSpan } from "confident-trace";
const generateResponse = span(
{ name: "generate_response", type: "llm", model: "gpt-4o", provider: "openai" },
async (prompt: string) => {
const response = await callLlm(prompt);
updateSpan({
inputTokenCount: response.usage.promptTokens,
outputTokenCount: response.usage.completionTokens,
});
return response.text;
},
);Token counts must be non-negative integers, and an explicit 0 is kept as 0 rather than treated as unset.
If you don't provide token counts and aren't using an integration, Confident AI will attempt to infer them by tokenizing the span's input and output text using the appropriate provider tokenizer. The table below summarizes each supported provider and its tokenization method.
| Provider | Tokenizer | Example Models | Token Counting Method |
|---|---|---|---|
| OpenAI | tiktoken | gpt-4o, gpt-4.1, o1, o3 | Client-side tokenization using model-specific encodings |
| Anthropic | @anthropic-ai/tokenizer | claude-3.5-sonnet, claude-3.7-sonnet, claude-4 | Claude-specific tokenization algorithm |
| Gemini API | gemini-2.0-flash, gemini-2.5-pro | Server-side token counting via API call |
See the OpenAI documentation, Anthropic documentation, or Google documentation for the most up-to-date pricing.
Track Token Usage Cost
Once token counts are available (either set manually, captured by an integration, or inferred automatically), Confident AI resolves the per-token cost using the following precedence:
- Per-token costs set in code — if you provide cost per input/output tokens directly via
span(...)orupdate_span()/updateSpan(), these always take priority. - Custom set model costs — if per-token costs aren't set in code, Confident AI uses the pricing configured in your Model Costs settings.
- Automatic price lookup — if no project-level costs are configured, Confident AI looks up the per-token pricing internally based on the
model. This is only available for OpenAI, Anthropic, and Gemini models.
If none of the above resolve a per-token cost, the cost for that side (input or output) is not logged.
Explicit Cost Setting
Set the per-token costs explicitly on the span alongside your token counts. This spares you for provider models not supported by automatic price lookup — self-hosted models, custom gateways, or negotiated enterprise pricing.
from confident_trace import span, update_span
@span(
type="llm", model="my-model", provider="custom",
cost_per_input_token=0.000001,
cost_per_output_token=0.000002,
)
def generate_response(prompt: str) -> str:
response = call_llm(prompt)
update_span(
input_token_count=response.usage.prompt_tokens,
output_token_count=response.usage.completion_tokens,
)
return response.textimport { span, updateSpan } from "confident-trace";
const generateResponse = span(
{
name: "generate_response", type: "llm", model: "my-model", provider: "custom",
costPerInputToken: 0.000001,
costPerOutputToken: 0.000002,
},
async (prompt: string) => {
const response = await callLlm(prompt);
updateSpan({
inputTokenCount: response.usage.promptTokens,
outputTokenCount: response.usage.completionTokens,
});
return response.text;
},
);You can pass the cost fields either on the span options (as above) or to update_span() / updateSpan() — whichever is more convenient. Either way they only apply to an LLM span; on any other span type the general fields are still applied but the LLM fields are skipped with a one-time warning.
Custom Price Lookup
If you provide token counts but don't set per-token costs in code, Confident AI will use the pricing you've configured in your project's Model Costs settings. This is useful when you want to manage pricing centrally without changing any code.
Model costs are matched against the model name on your LLM span using case-insensitive regex patterns, which match anywhere in the name unless anchored. For example:
^gpt-4o$— matches onlygpt-4ogpt-4.*— matchesgpt-4o,gpt-4o-mini,gpt-4-turbo, etc.claude-.*— matches all Claude model variants
You can optionally restrict a cost rule to a specific provider, and set input and output costs independently per million tokens. See the full Model Costs settings page for setup instructions.
Automatic Price Lookup
If you provide a supported model on your LLM span and neither SDK-level nor project-level costs are configured, Confident AI will automatically look up the per-token pricing and calculate the cost — no additional code needed.
from confident_trace import span, update_span
@span(type="llm", model="gpt-4o", provider="openai")
def generate_response(prompt: str) -> str:
output = call_llm(prompt)
update_span(input=prompt, output=output)
return outputimport { span, updateSpan } from "confident-trace";
const generateResponse = span(
{ name: "generate_response", type: "llm", model: "gpt-4o", provider: "openai" },
async (prompt: string) => {
const output = await callLlm(prompt);
updateSpan({ input: prompt, output });
return output;
},
);Cost on Traces
Cost on traces are automatically set by summing up the cost of all LLM spans in said trace. Similar to LLM spans, trace cost defaults to null values if no LLM spans have non-null values.
Next Steps
With cost tracking configured, continue setting up the rest of your instrumentation.
Set Input/Output
Override the default input and output on traces and spans for better visualization and evaluation.
Thread Traces
Group traces into threads to track multi-turn conversations and evaluate entire workflows.
Last updated on