LangGraph | Confident AI Docs

Overview

LangGraph is a framework for building reactive, multi-agent systems. Confident AI provides a CallbackHandler to trace and evaluate LangGraph agents.

Tracing Quickstart

Install Dependencies

Run the following command to install the required packages:

$ pip install -U deepeval langgraph langchain langchain-openai

Setup Confident AI Key

$ deepeval login

Configure LangGraph

Provide DeepEval’s CallbackHandler to your LangGraph agent’s invoke method.

main.py

1 from langgraph.prebuilt import create_react_agent
2 from langchain_openai import ChatOpenAI
3 
4 from deepeval.integrations.langchain import CallbackHandler
5 
6 def get_weather(city: str) -> str:
7     """Returns the weather in a city"""
8     return f"It's always sunny in {city}!"
9 
10 llm = ChatOpenAI(model="gpt-4o-mini")
11 
12 agent = create_react_agent(
13     model=llm,
14     tools=[get_weather],
15     prompt="You are a helpful assistant",
16 )
17 
18 result = agent.invoke(
19     input={"messages": [{"role": "user", "content": "what is the weather in sf"}]},
20     config={"callbacks": [CallbackHandler()]}
21 )

DeepEval’s CallbackHandler extends LangChain’s BaseCallbackHandler.

Run LangGraph

Invoke your agent by executing the script:

$ python main.py

You can directly view the traces on Confident AI by clicking on the link in the output printed in the console.

Advanced Features

Setting Trace Attributes

Confident AI’s LLM tracing advanced features provide teams with the ability to set certain attributes for each trace when invoking your LangChain application.

For example, thread_id and user_id are used to group related traces together, and are useful for chat apps, agents, or any multi-turn interactions. You can learn more about threads here.

You can set these attributes in the CallbackHandler when invoking your LangChain application.

main.py

1 result = agent_executor.invoke(
2     {"input": "What is 8 multiplied by 6?"},
3     config={
4         "callbacks": [CallbackHandler(thread_id="123")]
5     },
6 )

View Trace Attributes

name

str

The name of the trace. Learn more.

Logging prompts

If you are managing prompts on Confident AI and wish to log them, pass your Prompt object to the language model instance’s metadata parameter.

main.py

1 from langchain_openai import ChatOpenAI
2 from deepeval.prompt import Prompt
3 
4 prompt = Prompt(alias="<prompt-alias>")
5 prompt.pull(version="00.00.01")
6 
7 llm = ChatOpenAI(
8     model="gpt-4o-mini",
9     metadata={"prompt": prompt}
10 )

Logging prompts lets you attribute specific prompts to OpenAI Agent LLM spans. Be sure to pull the prompt before logging it, otherwise the prompt will not be visible on Confident AI.

Evals Usage

Online evals

If your LangChain application is in production, and you still want to run evaluations on your traces, use online evals. It lets you run evaluations on all incoming traces on Confident AI’s server.

Create metric collection

Create a metric collection on Confident AI with the metrics you wish to use to evaluate your LangGraph agent. Copy the name of the metric collection.

Create metric collection

The current LangChain integration supports metrics that only evaluate Input and Actual Output in addition to the Task Completion metric.

Run evals

Set the metric_collection name to evaluate various components of your LangChain application.

Agent Span

LLM Span

Tool Span

This is the top level component of your LangChain application. Also a very idle component to evaluate with the Task Completion metric.

main.py

1 from langchain_openai import ChatOpenAI
2 from langgraph.prebuilt import create_react_agent
3 
4 from deepeval.integrations.langchain import CallbackHandler
5 
6 def get_weather(city: str) -> str:
7     """Returns the weather in a city"""
8     return f"It's always sunny in {city}!"
9 
10 llm = ChatOpenAI(model="gpt-4o-mini")
11 
12 agent = create_react_agent(
13     model=llm,
14     tools=[get_weather],
15     prompt="You are a helpful assistant",
16 )
17 
18 result = agent.invoke(
19     input={"messages": [{"role": "user", "content": "what is the weather in sf"}]},
20     config={
21         "callbacks": [
22             CallbackHandler(metric_collection="task_completion")
23         ],
24     },
25 )

All incoming traces will now be evaluated using metrics from your metric collection.

View on Confident AI

You can view the evals on Confident AI by clicking on the link in the output printed in the console.