Shared evaluation dashboards
Every DeepEval test run is automatically synced to a shared dashboard. No more exporting CSVs or pasting results in Slack — your whole team sees the same metrics in real time.
DeepEval is the open-source LLM evaluation framework built and maintained by Confident AI. Run 50+ research-backed metrics in Python or TypeScript, then scale them org-wide with Confident AI's observability and collaboration.From the creators of DeepEval.
OTel-native, 30+ integrations, and SDKs in Python and TypeScript that feed straight into traces and evals.
Join the largest and fastest growing community on AI evaluation.
Checkout our FAQs below, or talk to a human. They won't hallucinate.
DeepEval is an open-source evaluation framework that lets you write and run LLM evaluation tests locally in Python or TypeScript. Confident AI is the cloud platform built on top of DeepEval that adds centralized test management, observability, collaboration, and analytics so teams can scale their evaluation workflows organization-wide.
Yes. Confident AI created DeepEval and maintains it today. Confident AI open-sourced DeepEval to give the community a best-in-class LLM evaluation framework, and extends it with the enterprise features teams need to operationalize evaluations at scale.
No. Confident AI builds both — DeepEval ships with native Confident AI support, and the platform works standalone if you never write a line of DeepEval. Confident AI is also a full LLM observability platform, so you can use it to trace, monitor, and evaluate your LLM applications in one place — no more siloing evals and tracing across different tools.
Yes. DeepEval is fully open-source under the Apache 2.0 license and free to use for any purpose. Confident AI offers a free tier as well, along with paid plans for teams that need advanced features like role-based access, custom dashboards, and dedicated support.
DeepEval ships with 50+ research-backed metrics including faithfulness, answer relevancy, contextual recall, contextual precision, hallucination, bias, toxicity, and more. You can also define fully custom metrics using Python or LLM-as-a-judge approaches.
Yes. While Confident AI has first-class support for DeepEval, it also integrates with other popular tools and frameworks through its REST API and SDKs, so you can centralize results regardless of how you run your evaluations.
Install DeepEval with pip install deepeval, write your first evaluation test, and optionally connect to Confident AI by running deepeval login. You can also sign up for Confident AI directly and start using the platform without DeepEval.