Confident AI Blog - Resources to help teams stay confident in AI
Launch Week 02 wrapped — explore all five launches

Stay Confident

Subscribe to our weekly newsletter to stay confident in the AI systems you build.

AI Agent Observability: Everything You Need to Know in 2026

AI Agent Observability: Everything You Need to Know in 2026

Everything you need to know about AI agent observability in 2026 — traces, spans, and threads; online and offline evals; production monitoring; and closing the feedback loop so failures never repeat.

Kritin Vongthongsri

Kritin Vongthongsri

Jun 25, 2026
.
13 min read
Human-in-the-Loop Workflows for AI Agent Evaluation: Complete Guide

Human-in-the-Loop Workflows for AI Agent Evaluation: Complete Guide

A practical guide to human-in-the-loop workflows for AI agent evaluation: how SMEs review AI agent failures, align automated metrics, and improve evaluation datasets.

Kritin Vongthongsri

Kritin Vongthongsri

Jun 13, 2026
.
10 min read
LLM Product Manager Workflows: A Complete Guide to AI Quality

LLM Product Manager Workflows: A Complete Guide to AI Quality

A practical guide to LLM product manager workflows, built around the two things PMs can finally do without waiting on engineering: build on the AI product by editing prompts, running evals, and comparing variants, and monitor quality with dashboards, signals, and shareable evidence.

Kritin Vongthongsri

Kritin Vongthongsri

Jun 13, 2026
.
9 min read
The Complete Guide to LLM Experimentation: Compare Prompts, Models, and Agents

The Complete Guide to LLM Experimentation: Compare Prompts, Models, and Agents

A practical guide to running LLM experiments across prompts, models, tools, datasets, metrics, production A/B tests, and human-in-the-loop feedback loops.

Kritin Vongthongsri

Kritin Vongthongsri

Jun 10, 2026
.
12 min read
LLM Evaluation for Startups: The Complete Guide

LLM Evaluation for Startups: The Complete Guide

A practical LLM evaluation guide for startups: build a small dataset, use the 2 + 3 metric rule, run CI/CD evals, and grow coverage from production signals and human review.

Kritin Vongthongsri

Kritin Vongthongsri

Jun 4, 2026
.
8 min read
Three Ways AI Systems Fail Even When Evals Pass

Three Ways AI Systems Fail Even When Evals Pass

AI systems can pass evals while still behaving incorrectly. This post explores three common failure modes that slip through output-based evaluation.

Brian Neville-O'Neill

Brian Neville-O'Neill

Apr 7, 2026
.
12 min
Your AI Agent Passed Evals. That’s the Problem.

Your AI Agent Passed Evals. That’s the Problem.

Passing evals doesn't mean your AI agent works — it means your tests missed how it fails. Why output-based evals create false confidence and what to measure instead.

Brian Neville-O'Neill

Brian Neville-O'Neill

Apr 6, 2026
.
4 min read
Multi-Turn LLM Evaluation in 2026: What You Need to Know

Multi-Turn LLM Evaluation in 2026: What You Need to Know

In this article, I'll break down multi-turn LLM evaluation — how it differs from single-turn, what metrics actually matter, and how to implement it.

Jeffrey Ip

Jeffrey Ip

Mar 22, 2026
.
14 min read
The Step-By-Step Guide to MCP Evaluation

The Step-By-Step Guide to MCP Evaluation

A step-by-step guide to MCP evaluation: how to test MCP-based LLM apps and agents, measure tool use and task completion, and catch failures with DeepEval.

Cale

Cale

Oct 25, 2025
.
9 min read
AI Agent Evaluation: Metrics, Traces, Human Review, and Workflows

AI Agent Evaluation: Metrics, Traces, Human Review, and Workflows

A practical guide to evaluating AI agents with LLM metrics and tracing—plus when human review matters, how it calibrates judges, and workflows that combine CI, sampling, and production signals.

Jeffrey Ip

Jeffrey Ip

Oct 7, 2025
.
20 min read