Open for registration
Webinar 01: From development to production.
Join Jeffrey and Tony for a live walkthrough of one AI agent: build a reusable eval harness, catch regressions before shipping, and keep evaluating on production traffic.
Jeffrey IpCEO and Co-founder @ Confident AI
Tony Cueva BravoFounding GTM @ Confident AI
AI agents act weird in production. Join us on September 30, 2026, 1pm - 1:45pm ET to learn how to:
- Score your app before it ships: upload a dataset of goldens, create a metric collection, and score your test cases against it, so every run gets the same yardstick.
- Catch regressions and ship the winner: run a regression test on a new version, compare the two runs side by side, and deploy the one that scores better.
- Keep scoring in production: set up tracing on the deployed app, create an evaluation rule, reuse the same metric collection, and score live traces as they come in.
- Close the loop: send the traces that failed back into your dataset as new goldens, so the next test run covers what production just taught you. Then run the cycle again.