October 2, 2026
- Human Feedback
- Evals
- Observability
- Integrations
- Security
Note-Worthy
TGIF! Thank god it's features, here's what we shipped this week:
Annotations got the glow-up this week. The annotations page now opens on scorecards and failure modes, so you see where reviewers agree, where they don't, and what keeps breaking before you've read a single item. Reviewers get hot keys to fly through the queue, queues can refill themselves on a schedule, alignment checks your setup before it runs, and error analysis finally comes with charts. Your human feedback has never been this easy to give or this easy to read. Consider it duly noted.

Added
- Revamped Annotation Analytics - The annotations page now leads with scorecards and failure modes. See at a glance how your reviewers are scoring, where they disagree, and which failures keep coming back, without exporting anything to a spreadsheet first. Less squinting at tables, more knowing the score.
- Auto-Annotate Hot Keys - Score, label, and jump to the next item without touching the mouse. Reviewers who live in the queue can now get through it at the speed they read. Your mouse gets the day off; your queue gets keyed up.
- Recurring Queue Ingestion Tasks - Queue ingestion tasks can now run on a schedule, so annotation queues fill themselves on whatever cadence you set. No more Monday-morning "did anyone load the queue?" Set it, forget it, review it. The queue goes on repeat.
- Validations for Metric Alignment - The alignment page now validates your setup before you run it, so you hear about a mismatched metric or missing label before the results come back looking odd. Measure twice, align once.
- Error Analysis Visualizations & Filters - Error analysis results now come as charts, and you can filter them down to the slice you care about. Spot the pattern, then narrow in on it. A picture is worth a thousand errors.
- EvalMode & Decision Models for Every Metric - Every metric on the platform now supports EvalMode, so you choose how it's judged:
LLM,HYBRID, orDECISION. Pick a decision model and let it answer bounded questions where a full LLM verdict is too much. Same metrics, more ways to decide. - Trajectory Custom Metrics - Build custom metrics that judge an agent's whole trajectory, not just the final answer. Grade the route as well as where it ended up. It's about the journey, and now you can score it.
- Span Filters for Traces, Signals & Dashboards - Filter by span on the traces page, on the signals page when it's nested under a trace, and on custom dashboards. Find the one span that mattered without scrolling through the ones that didn't. Span-tastic filtering.
- LiveKit Audio Recordings on Traces & Spans - Voice agents built on LiveKit can now attach audio recordings to traces and spans. Press play on the exact turn that went wrong instead of guessing from the transcript. Now we're hearing things, in a good way.
- Persona Metadata & Background Noise - Personas now take metadata and background noise, so simulated users show up with context and imperfect conditions, not studio-quiet ones. Real users are messy; now your simulations can be too. Make some noise.
- Findings Alerts for Regressions, Anomalies & Signal Spikes - Findings now send alerts when a metric regresses, something looks anomalous, or a signal spikes. Hear about it before your users do. We'll be alert so you don't have to be.
- Snowflake Export Integration - Export your Confident AI data straight into Snowflake and query it next to everything else in your warehouse. Your evals, out in the cold storage where they belong.
- Multiple Credentials per Project for Bedrock & Azure OpenAI - A single project can now hold several credentials for AWS Bedrock and Azure OpenAI. Split by region, account, or team without spinning up another project. More keys, same ring.
- OAuth 2.0, mTLS & Client Secret Auth for AI Connections - AI connections now support three new auth modes: OAuth 2.0, mTLS, and client secret. Connect endpoints behind real enterprise auth without a proxy in the middle. Whatever handshake your security team likes, we shake on it.
That's the drop for this week—see you next Friday.