•11 min read

LLM Observability and Eval-Driven Development with LangSmith: Tracing, Feedback & CI Gates

LLM Observability and Eval-Driven Development with LangSmith: Tracing, Feedback & CI Gates

This guide details the implementation of robust observability and evaluation-driven development pipelines for Large Language Model (LLM) applications using LangSmith. It covers distributed tracing for complex RAG pipelines and autonomous agents, ground-truth dataset creation, automated evaluation with exact-match and LLM-as-a-judge evaluators, and integration of evaluation gates into CI/CD workflows to prevent prompt regression.

Audio Briefing
0:00 / 0:00

1. Architectural Overview: LLM Observability and Evaluation

Production-grade LLM applications demand sophisticated observability and rigorous evaluation. Unlike traditional software, LLM behavior is non-deterministic, making traditional unit and integration testing insufficient. An effective LLM development lifecycle integrates continuous monitoring, feedback loops, and automated evaluation to ensure performance, reliability, and cost efficiency.

The core components of this architecture are:

  1. Distributed Tracing: Capturing the entire execution flow of an LLM application, from user input to final output, including all intermediate steps, API calls, and model interactions. This is crucial for debugging complex RAG chains and multi-agent systems.
  2. Feedback Mechanisms: Collecting human feedback on LLM outputs in production to identify failure modes and generate high-quality evaluation data.
  3. Ground Truth Dataset Curation: Building representative datasets with expected inputs and desired outputs for automated evaluation.
  4. Automated Evaluation: Running various evaluators (exact-match, semantic similarity, LLM-as-a-judge) against curated datasets to quantify performance.
  5. Evaluation-Driven Development (EDD): Iteratively improving LLM applications based on evaluation results, treating evaluation metrics as primary development targets.
  6. CI/CD Integration: Embedding evaluation gates into continuous integration pipelines to prevent regressions and enforce quality standards before deployment.

LangSmith serves as the central platform for orchestrating these components, providing a unified interface for tracing, dataset management, evaluation execution, and result visualization.

Advertisement

2. Setting Up LangSmith

Ensure you have a LangSmith account and API key.

import os
from dotenv import load_dotenv

load_dotenv()

os.environ["LANGCHAIN_TRACING_V2"] = "true"
os.environ["LANGCHAIN_API_KEY"] = os.getenv("LANGSMITH_API_KEY")
os.environ["LANGCHAIN_PROJECT"] = "locionic-llm-observability-guide" # Replace with your project name
os.environ["OPENAI_API_KEY"] = os.getenv("OPENAI_API_KEY") # Required for LLM-based evaluators and examples

3. Distributed Tracing for RAG and Agents

Distributed tracing in LangSmith provides a granular view of every operation within an LLM application. This is invaluable for debugging, performance optimization, and understanding complex interaction patterns.

3.1 Tracing a Simple RAG Chain

Consider a basic RAG pipeline: retrieve documents, then generate an answer based on the retrieved context.

from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from langchain_community.vectorstores import FAISS
from langchain.chains import create_retrieval_chain
from langchain.chains.combine_documents import create_stuff_documents_chain
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.documents import Document

# 1. Initialize Components
llm = ChatOpenAI(model="gpt-4o", temperature=0)
embeddings = OpenAIEmbeddings()

# 2. Create a dummy vector store
# In a real application, this would be populated from a knowledge base.
docs = [
    Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
    Document(page_content="Eiffel Tower is in Paris.", metadata={"source": "travel_guide"}),
    Document(page_content="The official language of France is French.", metadata={"source": "wiki"}),
    Document(page_content="Mount Everest is the highest mountain in the world.", metadata={"source": "geography"}),
]
vectorstore = FAISS.from_documents(docs, embeddings)
retriever = vectorstore.as_retriever()

# 3. Define the prompt for generation
system_prompt = (
    "You are an assistant for question-answering tasks. "
    "Use the following pieces of retrieved context to answer the question. "
    "If you don't know the answer, just say that you don't know. "
    "Keep the answer concise."
    "\n\n{context}"
)
prompt = ChatPromptTemplate.from_messages(
    [
        ("system", system_prompt),
        ("human", "{input}"),
    ]
)

# 4. Create the document combining chain
question_answer_chain = create_stuff_documents_chain(llm, prompt)

# 5. Create the full retrieval chain
rag_chain = create_retrieval_chain(retriever, question_answer_chain)

# 6. Invoke the chain and observe in LangSmith
print("Invoking RAG chain...")
response = rag_chain.invoke({"input": "What is the capital of France?"})
print(f"RAG Response: {response['answer']}")

response_no_answer = rag_chain.invoke({"input": "What is the capital of Germany?"})
print(f"RAG Response (no answer): {response_no_answer['answer']}")

After running this, navigate to your LangSmith project. You will see traces for each rag_chain.invoke() call. Each trace will show:

  • The initial rag_chain run.
  • A nested retriever run, showing the documents fetched.
  • A nested question_answer_chain run, showing the prompt constructed with context and the LLM call.
  • The final LLM response.

This hierarchical view is critical for identifying bottlenecks (e.g., slow retriever), prompt engineering issues (e.g., context not being used), or LLM hallucination.

3.2 Tracing an Autonomous Agent

Autonomous agents introduce more complexity due to their iterative decision-making and tool usage. LangSmith's tracing capabilities are even more vital here.

from langchain_openai import ChatOpenAI
from langchain import hub
from langchain.agents import AgentExecutor, create_openai_functions_agent
from langchain_core.tools import tool

# Define a custom tool
@tool
def get_current_weather(location: str) -> str:
    """Get the current weather in a given location."""
    if "london" in location.lower():
        return "It's cloudy with a chance of rain in London."
    elif "paris" in location.lower():
        return "It's sunny and warm in Paris."
    else:
        return "Weather data not available for this location."

tools = [get_current_weather]

# Get the agent prompt from LangChain Hub
prompt = hub.pull("hwchase17/openai-functions-agent")

# Initialize the LLM
llm = ChatOpenAI(model="gpt-4o", temperature=0)

# Create the agent
agent = create_openai_functions_agent(llm, tools, prompt)

# Create the agent executor
agent_executor = AgentExecutor(agent=agent, tools=tools, verbose=True)

# Invoke the agent
print("\nInvoking Agent Executor (London weather)...")
agent_executor.invoke({"input": "What's the weather like in London?"})

print("\nInvoking Agent Executor (Paris weather)...")
agent_executor.invoke({"input": "What's the weather like in Paris?"})

print("\nInvoking Agent Executor (Unknown location)...")
agent_executor.invoke({"input": "What's the weather like in Tokyo?"})

In LangSmith, the agent traces will show:

  • The main AgentExecutor run.
  • Nested Agent runs, showing the LLM's thought process and tool calls.
  • Individual Tool runs, indicating which tool was called and its output.
  • Subsequent Agent runs, showing the LLM processing the tool output.

This allows you to debug agent loops, incorrect tool selection, or issues with tool output interpretation.

4. Feedback Mechanisms and Dataset Curation

High-quality evaluation starts with high-quality data. LangSmith facilitates collecting human feedback and curating ground-truth datasets.

4.1 Collecting Human Feedback in Production

Integrate feedback mechanisms directly into your application's UI. When a user interacts with an LLM output, provide options to rate the response (e.g., "👍", "👎", "Correct", "Incorrect", "Harmful").

LangSmith provides an API to log feedback programmatically.

from langsmith import Client
from langsmith import RunStatus

client = Client()

# Simulate a previous run (e.g., from your RAG chain)
# In a real scenario, you'd get the run_id from the LangChain callback handler.
# For demonstration, let's create a dummy run.
dummy_run = client.create_run(
    project_name=os.environ["LANGCHAIN_PROJECT"],
    run_type="llm",
    name="dummy_llm_call_for_feedback",
    inputs={"prompt": "What is the capital of France?"},
    outputs={"completion": "Paris."},
    status=RunStatus.COMPLETED,
)
dummy_run_id = dummy_run.id

# Simulate user feedback
feedback_score = 1.0 # 1.0 for positive, 0.0 for negative
feedback_key = "user_score"
comment = "The answer was accurate and concise."

client.create_feedback(
    run_id=dummy_run_id,
    key=feedback_key,
    score=feedback_score,
    comment=comment,
    source_info={"user_id": "user_123", "session_id": "sess_abc"}, # Optional metadata
)
print(f"Feedback logged for run_id: {dummy_run_id}")

# Another example with negative feedback
dummy_run_negative = client.create_run(
    project_name=os.environ["LANGCHAIN_PROJECT"],
    run_type="llm",
    name="dummy_llm_call_for_feedback_negative",
    inputs={"prompt": "Who won the 2022 World Cup?"},
    outputs={"completion": "Brazil."}, # Incorrect
    status=RunStatus.COMPLETED,
)
dummy_run_negative_id = dummy_run_negative.id

client.create_feedback(
    run_id=dummy_run_negative_id,
    key="user_score",
    score=0.0,
    comment="Incorrect answer, Argentina won.",
)
print(f"Negative feedback logged for run_id: {dummy_run_negative_id}")

This feedback can then be used to filter traces, identify common failure modes, and prioritize improvements. More importantly, positive feedback can be used to automatically generate ground truth examples.

4.2 Creating Ground Truth Datasets

LangSmith allows you to create datasets from existing runs, upload CSV/JSON, or programmatically.

4.2.1 From Existing Runs (Feedback-Driven)

You can filter runs in LangSmith by feedback score and then export them to a dataset. For example, filter for runs with user_score:1.0 and then select "Create Dataset" from the UI. This is a powerful way to bootstrap evaluation datasets from production usage.

4.2.2 Programmatic Dataset Creation

For new features or specific test cases, you'll create datasets programmatically.

from langsmith import Client
from langsmith.schemas import Example

client = Client()

dataset_name = "RAG Question Answering Evaluation"
dataset_description = "Questions and answers for evaluating the RAG chain."

# Check if dataset exists, create if not
try:
    dataset = client.read_dataset(dataset_name=dataset_name)
    print(f"Dataset '{dataset_name}' already exists.")
except Exception:
    dataset = client.create_dataset(
        dataset_name=dataset_name,
        description=dataset_description,
        data_type="kv", # Key-value pairs
    )
    print(f"Dataset '{dataset_name}' created.")

# Define examples
examples = [
    {"input": "What is the capital of France?", "output": "Paris."},
    {"input": "Who painted the Mona Lisa?", "output": "Leonardo da Vinci."},
    {"input": "What is the highest mountain in the world?", "output": "Mount Everest."},
    {"input": "What is the largest ocean on Earth?", "output": "Pacific Ocean."},
    {"input": "What is the chemical symbol for water?", "output": "H2O."},
]

# Add examples to the dataset
for i, ex in enumerate(examples):
    # Check if example already exists to prevent duplicates on re-run
    # In a real scenario, you might have a more robust check or clear existing examples.
    try:
        client.create_example(
            dataset_id=dataset.id,
            inputs={"input": ex["input"]},
            outputs={"answer": ex["output"]}, # Match the output key of your chain
            metadata={"example_id": f"ex_{i}"}
        )
        print(f"Added example: {ex['input']}")
    except Exception as e:
        if "already exists" in str(e): # Basic check for existing example
            print(f"Example '{ex['input']}' already exists in dataset.")
        else:
            raise e

print(f"Dataset '{dataset_name}' populated with {len(examples)} examples.")

Key Considerations for Datasets:

  • Representativeness: Ensure your dataset covers common user queries, edge cases, and known failure modes.
  • Diversity: Include a variety of question types (factual, inferential, conversational).
  • Ground Truth Quality: The output (or reference_output) in your examples must be unequivocally correct.
  • Input/Output Keys: The keys in inputs and outputs dictionaries must match the expected input and output keys of the chain you are evaluating (e.g., input and answer for the RAG chain).
Advertisement

5. Automated Evaluation with LangSmith

Automated evaluation is the cornerstone of EDD. LangSmith provides various evaluators, from simple exact-match to sophisticated LLM-as-a-judge.

5.1 Running an Evaluation

To run an evaluation, you need:

  1. A dataset (created in Section 4.2).
  2. A "runnable" (your LLM chain or agent).
  3. One or more evaluators.
from langsmith import Client
from langsmith.evaluation import evaluate
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
from langchain.chains import create_retrieval_chain
from langchain.chains.combine_documents import create_stuff_documents_chain
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings
from langchain_core.documents import Document
import os

client = Client()

# Re-create the RAG chain for evaluation
llm_eval = ChatOpenAI(model="gpt-4o", temperature=0)
embeddings_eval = OpenAIEmbeddings()

docs_eval = [
    Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
    Document(page_content="Eiffel Tower is in Paris.", metadata={"source": "travel_guide"}),
    Document(page_content="The official language of France is French.", metadata={"source": "wiki"}),
    Document(page_content="Mount Everest is the highest mountain in the world.", metadata={"source": "geography"}),
    Document(page_content="Leonardo da Vinci painted the Mona Lisa.", metadata={"source": "art_history"}),
    Document(page_content="The Pacific Ocean is the largest ocean on Earth.", metadata={"source": "geography"}),
    Document(page_content="Water's chemical symbol is H2O.", metadata={"source": "chemistry"}),
]
vectorstore_eval = FAISS.from_documents(docs_eval, embeddings_eval)
retriever_eval = vectorstore_eval.as_retriever()

system_prompt_eval = (
    "You are an assistant for question-answering tasks. "
    "Use the following pieces of retrieved context to answer the question. "
    "If you don't know the answer, just say that you don't know. "
    "Keep the answer concise."
    "\n\n{context}"
)
prompt_eval = ChatPromptTemplate.from_messages(
    [
        ("system", system_prompt_eval),
        ("human", "{input}"),
    ]
)
question_answer_chain_eval = create_stuff_documents_chain(llm_eval, prompt_eval)
rag_chain_to_evaluate = create_retrieval_chain(retriever_eval, question_answer_chain_eval)

# Define evaluators
# 1. Exact Match Evaluator
# Checks if the predicted output exactly matches the reference output.
# Case-insensitive and whitespace-insensitive by default.
from langsmith.evaluation import ExactMatchEvaluator

exact_match_evaluator = ExactMatchEvaluator(
    criteria={"accuracy": "The predicted output should exactly match the reference output."}
)

# 2. LLM-as-a-Judge Evaluator
# Uses an LLM to assess the quality of the response based on custom criteria.
# Requires an LLM (e.g., OpenAI's gpt-4o).
from langsmith.evaluation import LangChainStringEvaluator

llm_judge_evaluator = LangChainStringEvaluator(
    eval_llm=llm_eval, # Use the same LLM or a more powerful one for evaluation
    config={
        "criteria": {
            "correctness": "Is the submission factually correct and directly answers the question?",
            "conciseness": "Is the submission concise and to the point, avoiding unnecessary verbosity?",
            "relevance": "Is the submission relevant to the question asked?",
        },
        "eval_chain_name": "llm_judge_rag_evaluator",
    },
    # You can also provide a custom prompt for the LLM judge
    # prompt="You are an expert evaluator. Assess the following submission..."
)

# Run the evaluation
print("Starting evaluation run...")
evaluation_results = evaluate(
    llm_or_chain=rag_chain_to_evaluate,
    data=dataset_name, # Name of the dataset created earlier
    evaluators=[
        exact_match_evaluator,
        llm_judge_evaluator,
        # You can add more evaluators here, e.g., "qa", "cot_qa", "embedding_distance"
    ],
    experiment_prefix="rag-chain-v1", # Name for this evaluation run
    metadata={"model_version": "gpt-4o", "retriever_config": "faiss_default"},
    max_concurrency=5, # Adjust based on API rate limits
)

print("\nEvaluation complete. Results available in LangSmith.")
# The evaluation_results object contains a summary and links to the full results in LangSmith.
# You can access metrics like:
# print(evaluation_results["results"][0].evaluator_info)

After the evaluation completes, navigate to the "Testing" section in your LangSmith project. You will see your rag-chain-v1 experiment. Click on it to view detailed results, including:

  • Aggregate scores for each evaluator.
  • Individual example results, showing inputs, reference outputs, predicted outputs, and evaluator feedback.
  • Traces for each evaluated run, allowing deep inspection of failures.

5.2 Types of Evaluators and Trade-offs

LangSmith offers a variety of built-in evaluators.

| Evaluator Type | Description

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement