•18 min read

LangGraph in Production: Human-in-the-Loop, Checkpointing & State Persistence

LangGraph in Production: Human-in-the-Loop, Checkpointing & State Persistence

LangGraph provides a robust framework for building stateful, multi-actor applications. Deploying these agentic workflows in production necessitates careful consideration of state management, fault tolerance, and human intervention. This guide details the implementation of production-grade LangGraph applications, focusing on state schemas, persistent checkpointing, human-in-the-loop (HITL) mechanisms, and dynamic subgraph management.

Audio Briefing
0:00 / 0:00

Core Concepts

LangGraph's power stems from its ability to model agentic workflows as directed acyclic graphs (DAGs) or cyclic graphs with state. Each node in the graph represents a step, and edges define transitions. The state object, passed between nodes, is central to maintaining context and enabling complex interactions.

State Schema Definition

A well-defined state schema is critical for maintainability, type safety, and data integrity. LangGraph leverages TypedDict and Pydantic for this purpose. Pydantic offers validation, serialization, and deserialization, making it ideal for production environments.

Consider an agent assisting with customer support, requiring information about the user's query, past interactions, and potential resolutions.

from typing import List, Literal, TypedDict, Optional
from langchain_core.messages import BaseMessage, HumanMessage, AIMessage
from pydantic import BaseModel, Field

# Define the state using TypedDict for basic structure
class AgentState(TypedDict):
    """
    Represents the state of our agentic workflow.
    """
    chat_history: List[BaseMessage]
    user_query: str
    tool_output: Optional[str]
    next_action: Optional[Literal["call_tool", "respond_to_user", "human_review"]]
    escalation_reason: Optional[str]
    # Pydantic models can be nested within TypedDict for richer validation
    customer_profile: Optional['CustomerProfile']

# Define a Pydantic model for a nested object within the state
class CustomerProfile(BaseModel):
    customer_id: str
    tier: Literal["bronze", "silver", "gold", "platinum"] = "bronze"
    recent_tickets: List[str] = Field(default_factory=list)
    is_vip: bool = False

# Example of how to use Pydantic for validation and default values
# This Pydantic model can be used to validate and parse the AgentState
class AgentStatePydantic(BaseModel):
    chat_history: List[BaseMessage]
    user_query: str
    tool_output: Optional[str] = None
    next_action: Optional[Literal["call_tool", "respond_to_user", "human_review"]] = None
    escalation_reason: Optional[str] = None
    customer_profile: Optional[CustomerProfile] = None

    class Config:
        arbitrary_types_allowed = True # Allow BaseMessage

In this example, AgentState defines the overall structure, while CustomerProfile is a Pydantic model providing structured validation for customer data. This separation enhances clarity and reusability.

Checkpointing and State Persistence

Production systems require state persistence to recover from failures, resume long-running processes, and enable debugging. LangGraph provides Checkpointer implementations for this.

MemorySaver (Development/Testing)

For local development or testing, MemorySaver is sufficient. It stores checkpoints in memory.

from langgraph.checkpoint.memory import MemorySaver
from langgraph.graph import StateGraph, END

# Define a simple graph for demonstration
class SimpleAgentState(TypedDict):
    value: int

def increment_node(state: SimpleAgentState):
    return {"value": state["value"] + 1}

builder = StateGraph(SimpleAgentState)
builder.add_node("increment", increment_node)
builder.set_entry_point("increment")
builder.add_edge("increment", END)

memory_checkpoint = MemorySaver()
graph_with_memory = builder.compile(checkpointer=memory_checkpoint)

# Run the graph
config = {"configurable": {"thread_id": "thread-1"}}
result = graph_with_memory.invoke({"value": 0}, config=config)
print(f"MemorySaver Result: {result}") # {'value': 1}

# Resume from checkpoint
result_resume = graph_with_memory.invoke({"value": 100}, config=config) # This will overwrite if not careful
# To truly resume, you'd typically load the state and then invoke
# For MemorySaver, subsequent invokes with the same thread_id will use the last state
result_resume_correct = graph_with_memory.invoke(None, config=config) # Invoking with None uses the last checkpointed state
print(f"MemorySaver Resumed Result: {result_resume_correct}") # {'value': 2}

PostgresSaver (Production)

For production, a persistent store like PostgreSQL is essential. PostgresSaver integrates with a PostgreSQL database to store and retrieve checkpoints.

First, ensure you have a PostgreSQL database running and the psycopg2-binary package installed (pip install psycopg2-binary).

import os
from langgraph.checkpoint.postgres import PostgresSaver
from langgraph.graph import StateGraph, END
from typing import TypedDict

# Database connection string
# In production, use environment variables or a secret management system
DATABASE_URL = os.getenv("DATABASE_URL", "postgresql://user:password@localhost:5432/langgraph_db")

# Ensure your database and table exist.
# A simple table creation script:
# CREATE TABLE IF NOT EXISTS checkpoints (
#     thread_id VARCHAR(255) PRIMARY NULL,
#     checkpoint JSONB NOT NULL,
#     metadata JSONB,
#     created_at TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP
# );

# Define a simple state
class PersistentAgentState(TypedDict):
    count: int
    messages: list[str]

# Define a node function
def process_step(state: PersistentAgentState):
    new_count = state["count"] + 1
    new_messages = state["messages"] + [f"Processed step {new_count}"]
    return {"count": new_count, "messages": new_messages}

# Build the graph
builder = StateGraph(PersistentAgentState)
builder.add_node("step_one", process_step)
builder.add_node("step_two", process_step)
builder.set_entry_point("step_one")
builder.add_edge("step_one", "step_two")
builder.add_edge("step_two", END)

# Initialize PostgresSaver
postgres_checkpoint = PostgresSaver(conn_string=DATABASE_URL)
graph_with_postgres = builder.compile(checkpointer=postgres_checkpoint)

# Example usage:
thread_id = "customer_interaction_123"
config = {"configurable": {"thread_id": thread_id}}

# Initial invocation
print(f"--- Initial Invocation for thread {thread_id} ---")
initial_state = {"count": 0, "messages": ["Start"]}
result_initial = graph_with_postgres.invoke(initial_state, config=config)
print(f"Result after first run: {result_initial}")
# Expected: {'count': 2, 'messages': ['Start', 'Processed step 1', 'Processed step 2']}

# Simulate a crash and resume
print(f"\n--- Resuming Invocation for thread {thread_id} ---")
# To resume, we invoke with None, which tells LangGraph to load the last checkpoint
result_resume = graph_with_postgres.invoke(None, config=config)
print(f"Result after resuming (should be same as initial if graph ended): {result_resume}")

# Let's modify the graph to show actual resumption
# Add another step and make it loop for demonstration
builder_loop = StateGraph(PersistentAgentState)
builder_loop.add_node("step_one", process_step)
builder_loop.add_node("step_two", process_step)
builder_loop.add_node("step_three", process_step) # New step
builder_loop.set_entry_point("step_one")
builder_loop.add_edge("step_one", "step_two")
builder_loop.add_edge("step_two", "step_three")
builder_loop.add_edge("step_three", END) # Changed to END for now

graph_with_postgres_loop = builder_loop.compile(checkpointer=postgres_checkpoint)

thread_id_loop = "customer_interaction_loop_456"
config_loop = {"configurable": {"thread_id": thread_id_loop}}

print(f"\n--- Looping Invocation for thread {thread_id_loop} ---")
# First run, will go through all steps
result_loop_1 = graph_with_postgres_loop.invoke({"count": 0, "messages": ["Loop Start"]}, config=config_loop)
print(f"Result after first loop run: {result_loop_1}")
# Expected: {'count': 3, 'messages': ['Loop Start', 'Processed step 1', 'Processed step 2', 'Processed step 3']}

# Now, let's simulate a partial run and then resume
# We need to modify the graph to *not* end immediately
builder_partial = StateGraph(PersistentAgentState)
builder_partial.add_node("step_A", process_step)
builder_partial.add_node("step_B", process_step)
builder_partial.add_node("step_C", process_step)
builder_partial.set_entry_point("step_A")
builder_partial.add_edge("step_A", "step_B")
# Intentionally stop before C to demonstrate resumption
# We'll use a conditional edge to simulate a breakpoint or partial run
def should_continue(state: PersistentAgentState):
    if state["count"] < 2: # Stop after step_B
        return "continue"
    return "end"

builder_partial.add_conditional_edges(
    "step_B",
    should_continue,
    {"continue": "step_C", "end": END}
)
builder_partial.add_edge("step_C", END)

graph_with_postgres_partial = builder_partial.compile(checkpointer=postgres_checkpoint)

thread_id_partial = "partial_run_789"
config_partial = {"configurable": {"thread_id": thread_id_partial}}

print(f"\n--- Partial Invocation for thread {thread_id_partial} ---")
# Run partially, it should stop after step_B
result_partial_1 = graph_with_postgres_partial.invoke({"count": 0, "messages": ["Partial Start"]}, config=config_partial)
print(f"Result after partial run 1 (should stop at count 2): {result_partial_1}")
# Expected: {'count': 2, 'messages': ['Partial Start', 'Processed step 1', 'Processed step 2']}

# Resume the partial run. It should pick up from where it left off (after step_B)
print(f"\n--- Resuming Partial Invocation for thread {thread_id_partial} ---")
result_partial_2 = graph_with_postgres_partial.invoke(None, config=config_partial)
print(f"Result after resuming partial run (should complete step_C): {result_partial_2}")
# Expected: {'count': 3, 'messages': ['Partial Start', 'Processed step 1', 'Processed step 2', 'Processed step 3']}

Human-in-the-Loop (HITL)

HITL is crucial for complex, high-stakes, or ambiguous workflows. LangGraph facilitates this via the interrupt mechanism. An agent can pause execution, await human input, and then resume.

from langgraph.graph import StateGraph, END
from typing import TypedDict, Literal, Optional
from langchain_core.messages import BaseMessage, HumanMessage, AIMessage

class HumanReviewState(TypedDict):
    chat_history: List[BaseMessage]
    user_query: str
    agent_response: Optional[str]
    human_feedback: Optional[str]
    review_needed: Literal["yes", "no"]

def agent_decides_review(state: HumanReviewState):
    # Simulate agent logic to decide if human review is needed
    if "sensitive" in state["user_query"].lower():
        return {"review_needed": "yes", "agent_response": "I've drafted a response, but it might be sensitive. Awaiting human review."}
    return {"review_needed": "no", "agent_response": "Here's my direct response."}

def generate_response(state: HumanReviewState):
    # Simulate generating a response
    if state["review_needed"] == "no":
        return {"agent_response": f"Agent's direct response to: {state['user_query']}"}
    # If review is needed, the agent_response would have been set by agent_decides_review
    return state

def human_review_node(state: HumanReviewState):
    # This node is where the human would provide feedback.
    # In a real system, this would be an API endpoint or UI interaction.
    # For demonstration, we'll assume feedback is provided externally.
    print(f"\n--- HUMAN REVIEW REQUIRED ---")
    print(f"User Query: {state['user_query']}")
    print(f"Agent's Draft/Decision: {state['agent_response']}")
    print(f"Please provide feedback for thread_id: {config['configurable']['thread_id']}")
    # The graph will be interrupted here.
    # The human would then call the graph with updated state via `update_state` or `invoke` with feedback.
    return state # State remains unchanged until human input

def process_human_feedback(state: HumanReviewState):
    if state["human_feedback"]:
        final_response = f"Agent incorporated human feedback: '{state['human_feedback']}'. Final response: {state['agent_response']}"
        return {"agent_response": final_response, "review_needed": "no"}
    return state

# Build the graph
builder = StateGraph(HumanReviewState)
builder.add_node("decide_review", agent_decides_review)
builder.add_node("generate_response", generate_response)
builder.add_node("human_review", human_review_node)
builder.add_node("process_feedback", process_human_feedback)

builder.set_entry_point("decide_review")

# Conditional edge for review
builder.add_conditional_edges(
    "decide_review",
    lambda state: state["review_needed"],
    {
        "yes": "human_review",
        "no": "generate_response",
    }
)

builder.add_edge("generate_response", END)
builder.add_edge("human_review", "process_feedback")
builder.add_edge("process_feedback", END)

# Compile with a checkpointer and interrupt_before
# Interrupt before 'human_review' node
memory_checkpoint = MemorySaver()
graph_with_hitl = builder.compile(
    checkpointer=memory_checkpoint,
    interrupt_before=["human_review"] # Interrupt before this node
)

# --- Scenario 1: No human review needed ---
thread_id_no_review = "hitl_thread_1"
config = {"configurable": {"thread_id": thread_id_no_review}}
print(f"\n--- Running Scenario 1: No Human Review Needed ({thread_id_no_review}) ---")
result_no_review = graph_with_hitl.invoke(
    {"chat_history": [], "user_query": "What is your return policy?", "review_needed": "no"},
    config=config
)
print(f"Final state (no review): {result_no_review}")
# Expected: {'chat_history': [], 'user_query': 'What is your return policy?', 'agent_response': "Agent's direct response to: What is your return policy?", 'human_feedback': None, 'review_needed': 'no'}

# --- Scenario 2: Human review needed ---
thread_id_review = "hitl_thread_2"
config = {"configurable": {"thread_id": thread_id_review}}
print(f"\n--- Running Scenario 2: Human Review Needed ({thread_id_review}) ---")
initial_state_review = {"chat_history": [], "user_query": "I have a sensitive issue with my account.", "review_needed": "yes"}

# First invocation: will run up to 'human_review' and interrupt
print("First invocation (interrupting at human_review)...")
try:
    # LangGraph raises a StopIteration when interrupted
    graph_with_hitl.invoke(initial_state_review, config=config)
except StopIteration as e:
    print(f"Graph interrupted at: {e.args[0]['configurable']['checkpoint_id']}")
    # The state at interruption can be retrieved from the checkpointer
    interrupted_state = graph_with_hitl.get_state(config)
    print(f"State at interruption: {interrupted_state.current}")

    # Simulate human providing feedback
    human_provided_feedback = "Ensure empathy and offer a direct contact number."
    print(f"\n--- Human provides feedback: '{human_provided_feedback}' ---")

    # Resume the graph with human feedback
    print("Resuming graph with human feedback...")
    final_result_review = graph_with_hitl.invoke(
        {"human_feedback": human_provided_feedback}, # Only provide the delta
        config=config
    )
    print(f"Final state after human review: {final_result_review}")
    # Expected: {'chat_history': [], 'user_query': 'I have a sensitive issue with my account.', 'agent_response': "Agent incorporated human feedback: 'Ensure empathy and offer a direct contact number.'. Final response: I've drafted a response, but it might be sensitive. Awaiting human review.", 'human_feedback': 'Ensure empathy and offer a direct contact number.', 'review_needed': 'no'}

The interrupt_before parameter in compile tells LangGraph to pause execution before entering the specified node. The StopIteration exception signals this pause. The human can then provide input, and the graph can be resumed by invoking it again with the updated state.

State Rollback

Checkpointing enables state rollback. If an agent makes an undesirable decision or enters an erroneous state, a human operator or automated system can revert to a previous checkpoint.

# Assuming graph_with_postgres_partial from above is compiled with PostgresSaver
# thread_id_partial = "partial_run_789"

# Get the history of checkpoints for a thread
history = graph_with_postgres_partial.get_state_history(config_partial)
print(f"\n--- Checkpoint History for {thread_id_partial} ---")
for i, checkpoint in enumerate(history):
    print(f"Checkpoint {i}: {checkpoint.config['configurable']['checkpoint_id']} - State: {checkpoint.state.current}")

# Let's say we want to rollback to the state before the last step
# The history is ordered from oldest to newest.
# To rollback, we need the ID of the checkpoint we want to restore.
# For demonstration, let's assume we want to go back to the state after 'step_B'
# which was the state before the final 'step_C' in the 'partial_run_789' example.
# This would be the second to last checkpoint in the history.

# Get the checkpoint ID of the state we want to restore
# In a real system, you'd have a UI or API to select this.
# For this example, let's manually pick the checkpoint ID from the history.
# Assuming the history has at least 2 checkpoints for 'partial_run_789'
if len(history) >= 2:
    checkpoint_to_restore_id = history[-2].config['configurable']['checkpoint_id']
    print(f"\n--- Rolling back to checkpoint ID: {checkpoint_to_restore_id} ---")

    # To rollback, we invoke with a specific checkpoint_id in the config
    rollback_config = {
        "configurable": {
            "thread_id": thread_id_partial,
            "checkpoint_id": checkpoint_to_restore_id
        }
    }
    # Invoking with None will load the specified checkpoint
    rolled_back_state = graph_with_postgres_partial.invoke(None, config=rollback_config)
    print(f"State after rollback: {rolled_back_state}")
    # Expected: {'count': 2, 'messages': ['Partial Start', 'Processed step 1', 'Processed step 2']}

    # Now, if we run the graph again without a specific checkpoint_id,
    # it will continue from the rolled-back state.
    print(f"\n--- Resuming after rollback ---")
    result_after_rollback = graph_with_postgres_partial.invoke(None, config=config_partial)
    print(f"Result after resuming from rolled-back state: {result_after_rollback}")
    # Expected: {'count': 3, 'messages': ['Partial Start', 'Processed step 1', 'Processed step 2', 'Processed step 3']}
else:
    print("Not enough checkpoints in history to demonstrate rollback for 'partial_run_789'.")

Dynamic Subgraphs

LangGraph allows for dynamic graph construction or modification, enabling adaptive workflows. This is particularly useful for agents that need to dynamically select tools or sub-agents based on context. While full dynamic graph structure changes at runtime are complex, dynamic invocation of subgraphs or conditional routing is common.

from langgraph.graph import StateGraph, END, START
from typing import TypedDict, Literal, Optional
from langchain_core.messages import BaseMessage, HumanMessage, AIMessage

class DynamicAgentState(TypedDict):
    chat_history: List[BaseMessage]
    current_task: Optional[str]
    tool_output: Optional[str]
    next_step: Literal["plan", "execute_search", "execute_calculator", "respond", "end"]

# Define some dummy tool nodes
def search_tool_node(state: DynamicAgentState):
    query = state["current_task"]
    print(f"Executing search for: {query}")
    # Simulate API call
    return {"tool_output": f"Search results for '{query}': Found 10 relevant documents."}

def calculator_tool_node(state: DynamicAgentState):
    expression = state["current_task"]
    print(f"Executing calculator for: {expression}")
    # Simulate API call
    try:
        result = eval(expression) # DANGER: Never use eval with untrusted input in production!
        return {"tool_output": f"Calculator result for '{expression}': {result}"}
    except Exception as e:
        return {"tool_output": f"Calculator error: {e}"}

def planner_node(state: DynamicAgentState):
    user_query = state["chat_history"][-1].content if state["chat_history"] else ""
    if "calculate" in user_query.lower() or "math" in user_query.lower():
        return {"current_task": "2 + 2 * 3", "next_step": "execute_calculator"}
    elif "search" in user_query.lower() or "find" in user_query.lower():
        return {"current_task": "latest AI news", "next_step": "execute_search"}
    else:
        return {"next_step": "respond"}

def responder_node(state: DynamicAgentState):
    response = f"Understood. {state.get('tool_output', 'No specific tool output.')} How else can I help?"
    return {"chat_history": state["chat_history"] + [AIMessage(content=response)], "next_step": "end"}

# Build the graph
builder = StateGraph(DynamicAgentState)
builder.add_node("planner", planner_node)
builder.add_node("execute_search", search_tool_node)
builder.add_node("execute_calculator", calculator_tool_node)
builder.add_node("responder", responder_node)

builder.set_entry_point("planner")

# Conditional routing based on planner's decision
builder.add_conditional_edges(
    "planner",
    lambda state: state["next_step"],
    {
        "execute_search": "execute_search",
        "execute_calculator": "execute_calculator",
        "respond": "responder",
    }
)

# After tool execution, go to responder
builder.add_edge("execute_search", "responder")
builder.add_edge("execute_calculator", "responder")
builder.add_edge("responder", END)

graph_dynamic = builder.compile(checkpointer=MemorySaver())

# --- Scenario 1: Search query ---
thread_id_search = "dynamic_thread_search"
config_search = {"configurable": {"thread_id": thread_id_search}}
print(f"\n--- Running Dynamic Scenario 1: Search ({thread_id_search}) ---")
result_search = graph_dynamic.invoke(
    {"chat_history": [HumanMessage(content="Please search for the latest AI news.")], "next_step": "plan"},
    config=config_search
)
print(f"Final state (search): {result_search}")
# Expected: tool_output with search results, chat_history updated.

# --- Scenario 2: Calculator query ---
thread_id_calc = "dynamic_thread_calc"
config_calc = {"configurable": {"thread_id": thread_id_calc}}
print(f"\n--- Running Dynamic Scenario 2: Calculator ({thread_id_calc}) ---")
result_calc = graph_dynamic.invoke(
    {"chat_history": [HumanMessage(content="Can you calculate 5 * 8 + 10?")], "next_step": "plan"},
    config=config_calc
)
print(f"Final state (calculator): {result_calc}")
# Expected: tool_output with calculation result, chat_history updated.

# --- Scenario 3: Direct response ---
thread_id_direct = "dynamic_thread_direct"
config_direct = {"configurable": {"thread_id": thread_direct}}
print(f"\n--- Running Dynamic Scenario 3: Direct Response ({thread_id_direct}) ---")
result_direct = graph_dynamic.invoke(
    {"chat_history": [HumanMessage(content="Hello there!")], "next_step": "plan"},
    config=config_direct
)
print(f"Final state (direct): {result_direct}")
# Expected: tool_output None, chat_history updated with direct response.

This example demonstrates dynamic routing based on the planner_node's output, effectively creating a dynamic subgraph execution path.

Advertisement

Architectural Comparison: Checkpointers

FeatureMemorySaverPostgresSaver
PersistenceNone (in-memory)Durable (PostgreSQL)
ScalabilityLow (single process)High (database-backed)
ConcurrencyLimitedHigh (database handles)
Use CaseDev, testing, ephemeral tasksProduction, long-running workflows, fault tolerance
SetupTrivialRequires PostgreSQL instance & schema
Data IntegrityLow (volatile)High (ACID properties)
RollbackPossible (if history kept)Robust (via checkpoint IDs)

Production Gotchas & Troubleshooting

  1. Schema Evolution: Changing your TypedDict or Pydantic state schema in production requires a migration strategy for existing checkpoints.
    • Gotcha: Adding a non-optional field without a default value will break loading old checkpoints.
    • Fix: Always add new fields as Optional or with a default value. For mandatory new fields, write a data migration script to update old checkpoints in the database.
  2. Database Connection Management: PostgresSaver creates a new connection for each operation. For high-throughput applications, this can lead to connection exhaustion.
    • Gotcha: "Too many connections" errors in PostgreSQL logs.
    • Fix: Implement a connection pool (e.g., using SQLAlchemy's create_engine with pool_size and max_overflow) and pass the connection object to PostgresSaver or manage it externally. Ensure connections are properly closed.
  3. Large State Objects: Storing very large objects (e.g., extensive chat histories, embedded documents) directly in the state can degrade performance and exceed database limits.
    • Gotcha: Slow checkpointing, large database storage, potential JSONB size limits.
    • Fix: Store references (e.g., IDs) to large objects in the state, and retrieve the actual data from a dedicated document store or object storage (S3, GCS) when needed by a node.
  4. Concurrency Issues with thread_id: If multiple instances or requests try to update the same thread_id concurrently without proper locking, race conditions can occur.
    • Gotcha: Inconsistent state, lost updates.
    • Fix: LangGraph's checkpointer handles basic atomicity for a single invoke call. For complex concurrent updates to the same thread, ensure your application layer implements appropriate locking or queueing mechanisms. Consider using a message queue (Kafka, RabbitMQ) to serialize updates for a given thread_id.
  5. interrupt_before in Production: While powerful, relying solely on StopIteration for HITL in a web service context can be tricky.
    • Gotcha: StopIteration is an exception, not a normal return. Handling it gracefully in an API endpoint requires specific error handling.
    • Fix: Design your API to catch StopIteration, extract the checkpoint_id and current state, and return it to the client. The client then presents this to the human and sends back the updated state to resume. Ensure the checkpoint_id is used for resumption.
  6. Security of eval() in Dynamic Subgraphs: As noted in the dynamic subgraph example, using eval() for dynamic code execution is a severe security vulnerability.
    • Gotcha: Remote code execution.
    • Fix: NEVER use eval() with untrusted input. For dynamic tool selection, use a predefined set of tools and map string names to function calls, or use a secure sandbox environment if arbitrary code execution is truly required (which is rare).

Frequently Asked Questions

  1. How do I handle authentication and authorization for human-in-the-loop steps? Implement an API layer that wraps your LangGraph invocation. When an interrupt_before node is hit, your API should return the current state and a checkpoint_id. The UI/client then presents this to an authenticated user. When the user provides feedback, the API receives it, authenticates the user, authorizes them to act on that specific thread_id, and then invokes LangGraph with the updated state and checkpoint_id to resume.
  2. Can I use a different database for checkpointing, like MongoDB or Redis? LangGraph provides MemorySaver and PostgresSaver out-of-the-box. For other databases, you would need to implement a custom BaseCheckpointSaver class. This involves implementing methods like get, put, and list. For MongoDB, you might use pymongo and store checkpoints as documents. For Redis, you could serialize the state to JSON and store it as a hash or string.
  3. What's the best way to monitor LangGraph workflows in production? Integrate with observability tools.
    • Logging: Use structured logging within your nodes to capture key events, state changes, and tool calls.
    • Tracing: LangChain integrates with LangSmith for detailed tracing of agent steps, LLM calls, and tool invocations. This is invaluable for debugging and understanding complex workflows.
    • Metrics: Emit custom metrics (e.g., node execution times, number of human interventions, error rates) to your Prometheus/Grafana or Datadog setup.
  4. How can I version my LangGraph workflows? Treat your graph definition as code and manage it in version control (Git). When deploying, ensure the deployed graph version is compatible with existing checkpoint schemas. For major schema changes, deploy a new version of the graph and potentially migrate old threads to the new schema or run old threads with the old graph definition. Consider using a versioning scheme for your thread_ids or storing the graph version within the checkpoint metadata.
  5. My LangGraph application is slow. How do I debug performance?
    • LLM Calls: The primary bottleneck is often LLM inference time. Optimize prompts, use faster models, or implement caching for common LLM calls.
    • Tool Calls: Profile your tool functions. Are external APIs slow? Are database queries optimized?
    • State Size: As mentioned, large state objects can slow down checkpointing.
    • Graph Complexity: Very complex graphs with many nodes and conditional edges can have overhead. Simplify where possible.
    • Checkpointer: Ensure your database (for PostgresSaver) is performant and properly indexed.
    • Tracing: Use LangSmith or similar tools to identify the slowest parts of your graph execution.
Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement