•16 min read

LangSmithによるLLMの可観測性と評価駆動開発: トレース、フィードバック、CIゲート

LangSmithによるLLMの可観測性と評価駆動開発: トレース、フィードバック、CIゲート

このガイドでは、LangSmith を使用した大規模言語モデル(LLM)アプリケーション向けの堅牢な可観測性(observability)と評価駆動型開発(evaluation-driven development)パイプラインの実装について詳しく説明します。複雑な RAG パイプラインと自律エージェントのための分散トレーシング、グラウンドトゥルースデータセットの作成、完全一致(exact-match)および LLM-as-a-judge 評価器による自動評価、そしてプロンプトの回帰を防ぐための CI/CD ワークフローへの評価ゲートの統合について解説します。

Audio Briefing
0:00 / 0:00

1. アーキテクチャの概要:LLM の可観測性と評価

本番環境レベルの LLM アプリケーションには、高度な可観測性と厳密な評価が求められます。従来のソフトウェアとは異なり、LLM の動作は非決定論的であるため、従来の単体テストや統合テストでは不十分です。効果的な LLM 開発ライフサイクルでは、継続的な監視、フィードバックループ、自動評価を統合し、パフォーマンス、信頼性、コスト効率を確保します。

このアーキテクチャの主要コンポーネントは以下の通りです。

  1. 分散トレーシング(Distributed Tracing): ユーザー入力から最終出力まで、すべての中間ステップ、API 呼び出し、モデルインタラクションを含む LLM アプリケーションの実行フロー全体を捕捉します。これは、複雑な RAG チェーンやマルチエージェントシステムのデバッグに不可欠です。
  2. フィードバックメカニズム(Feedback Mechanisms): 本番環境での LLM 出力に関する人間からのフィードバックを収集し、失敗モードを特定し、高品質な評価データを生成します。
  3. グラウンドトゥルースデータセットのキュレーション(Ground Truth Dataset Curation): 自動評価のために、期待される入力と望ましい出力を含む代表的なデータセットを構築します。
  4. 自動評価(Automated Evaluation): キュレーションされたデータセットに対して、さまざまな評価器(完全一致、意味的類似性、LLM-as-a-judge)を実行し、パフォーマンスを定量化します。
  5. 評価駆動型開発(Evaluation-Driven Development: EDD): 評価結果に基づいて LLM アプリケーションを反復的に改善し、評価指標を主要な開発目標として扱います。
  6. CI/CD 統合(CI/CD Integration): 継続的インテグレーションパイプラインに評価ゲートを組み込み、デプロイ前に回帰を防ぎ、品質基準を強制します。

LangSmith は、これらのコンポーネントをオーケストレーションするための中央プラットフォームとして機能し、トレーシング、データセット管理、評価実行、結果の可視化のための統合インターフェースを提供します。

Advertisement

2. LangSmith のセットアップ

LangSmith アカウントと API キーがあることを確認してください。

import os
from dotenv import load_dotenv

load_dotenv()

os.environ["LANGCHAIN_TRACING_V2"] = "true"
os.environ["LANGCHAIN_API_KEY"] = os.getenv("LANGSMITH_API_KEY")
os.environ["LANGCHAIN_PROJECT"] = "locionic-llm-observability-guide" # Replace with your project name
os.environ["OPENAI_API_KEY"] = os.getenv("OPENAI_API_KEY") # Required for LLM-based evaluators and examples

3. RAG とエージェントのための分散トレーシング

LangSmith の分散トレーシングは、LLM アプリケーション内のすべての操作を詳細に可視化します。これは、デバッグ、パフォーマンス最適化、複雑なインタラクションパターンの理解に非常に役立ちます。

3.1 シンプルな RAG チェーンのトレーシング

基本的な RAG パイプラインを考えてみましょう。ドキュメントを検索し、検索されたコンテキストに基づいて回答を生成します。

from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from langchain_community.vectorstores import FAISS
from langchain.chains import create_retrieval_chain
from langchain.chains.combine_documents import create_stuff_documents_chain
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.documents import Document

# 1. Initialize Components
llm = ChatOpenAI(model="gpt-4o", temperature=0)
embeddings = OpenAIEmbeddings()

# 2. Create a dummy vector store
# In a real application, this would be populated from a knowledge base.
docs = [
    Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
    Document(page_content="Eiffel Tower is in Paris.", metadata={"source": "travel_guide"}),
    Document(page_content="The official language of France is French.", metadata={"source": "wiki"}),
    Document(page_content="Mount Everest is the highest mountain in the world.", metadata={"source": "geography"}),
]
vectorstore = FAISS.from_documents(docs, embeddings)
retriever = vectorstore.as_retriever()

# 3. Define the prompt for generation
system_prompt = (
    "You are an assistant for question-answering tasks. "
    "Use the following pieces of retrieved context to answer the question. "
    "If you don't know the answer, just say that you don't know. "
    "Keep the answer concise."
    "\n\n{context}"
)
prompt = ChatPromptTemplate.from_messages(
    [
        ("system", system_prompt),
        ("human", "{input}"),
    ]
)

# 4. Create the document combining chain
question_answer_chain = create_stuff_documents_chain(llm, prompt)

# 5. Create the full retrieval chain
rag_chain = create_retrieval_chain(retriever, question_answer_chain)

# 6. Invoke the chain and observe in LangSmith
print("Invoking RAG chain...")
response = rag_chain.invoke({"input": "What is the capital of France?"})
print(f"RAG Response: {response['answer']}")

response_no_answer = rag_chain.invoke({"input": "What is the capital of Germany?"})
print(f"RAG Response (no answer): {response_no_answer['answer']}")

これを実行した後、LangSmith プロジェクトに移動します。各 rag_chain.invoke() 呼び出しのトレースが表示されます。各トレースには以下が含まれます。

  • 最初の rag_chain 実行。
  • ネストされた retriever 実行(フェッチされたドキュメントを表示)。
  • ネストされた question_answer_chain 実行(コンテキストと LLM 呼び出しで構築されたプロンプトを表示)。
  • 最終的な LLM レスポンス。

この階層的なビューは、ボトルネック(例:リトリーバーの遅延)、プロンプトエンジニアリングの問題(例:コンテキストが使用されていない)、または LLM のハルシネーションを特定するために重要です。

3.2 自律エージェントのトレーシング

自律エージェントは、反復的な意思決定とツール使用のため、より複雑になります。LangSmith のトレーシング機能は、ここではさらに不可欠です。

from langchain_openai import ChatOpenAI
from langchain import hub
from langchain.agents import AgentExecutor, create_openai_functions_agent
from langchain_core.tools import tool

# Define a custom tool
@tool
def get_current_weather(location: str) -> str:
    """Get the current weather in a given location."""
    if "london" in location.lower():
        return "It's cloudy with a chance of rain in London."
    elif "paris" in location.lower():
        return "It's sunny and warm in Paris."
    else:
        return "Weather data not available for this location."

tools = [get_current_weather]

# Get the agent prompt from LangChain Hub
prompt = hub.pull("hwchase17/openai-functions-agent")

# Initialize the LLM
llm = ChatOpenAI(model="gpt-4o", temperature=0)

# Create the agent
agent = create_openai_functions_agent(llm, tools, prompt)

# Create the agent executor
agent_executor = AgentExecutor(agent=agent, tools=tools, verbose=True)

# Invoke the agent
print("\nInvoking Agent Executor (London weather)...")
agent_executor.invoke({"input": "What's the weather like in London?"})

print("\nInvoking Agent Executor (Paris weather)...")
agent_executor.invoke({"input": "What's the weather like in Paris?"})

print("\nInvoking Agent Executor (Unknown location)...")
agent_executor.invoke({"input": "What's the weather like in Tokyo?"})

LangSmith では、エージェントのトレースに以下が表示されます。

  • メインの AgentExecutor 実行。
  • ネストされた Agent 実行(LLM の思考プロセスとツール呼び出しを表示)。
  • 個々の Tool 実行(どのツールが呼び出され、その出力が表示)。
  • 後続の Agent 実行(LLM がツール出力を処理していることを表示)。

これにより、エージェントのループ、誤ったツール選択、またはツール出力の解釈に関する問題をデバッグできます。

4. フィードバックメカニズムとデータセットのキュレーション

高品質な評価は、高品質なデータから始まります。LangSmith は、人間からのフィードバックの収集とグラウンドトゥルースデータセットのキュレーションを容易にします。

4.1 本番環境での人間からのフィードバックの収集

アプリケーションの UI にフィードバックメカニズムを直接統合します。ユーザーが LLM の出力と対話する際に、応答を評価するオプション(例:「👍」、「👎」、「正しい」、「間違っている」、「有害」)を提供します。

LangSmith は、フィードバックをプログラムでログに記録するための API を提供します。

from langsmith import Client
from langsmith import RunStatus

client = Client()

# Simulate a previous run (e.g., from your RAG chain)
# In a real scenario, you'd get the run_id from the LangChain callback handler.
# For demonstration, let's create a dummy run.
dummy_run = client.create_run(
    project_name=os.environ["LANGCHAIN_PROJECT"],
    run_type="llm",
    name="dummy_llm_call_for_feedback",
    inputs={"prompt": "What is the capital of France?"},
    outputs={"completion": "Paris."},
    status=RunStatus.COMPLETED,
)
dummy_run_id = dummy_run.id

# Simulate user feedback
feedback_score = 1.0 # 1.0 for positive, 0.0 for negative
feedback_key = "user_score"
comment = "The answer was accurate and concise."

client.create_feedback(
    run_id=dummy_run_id,
    key=feedback_key,
    score=feedback_score,
    comment=comment,
    source_info={"user_id": "user_123", "session_id": "sess_abc"}, # Optional metadata
)
print(f"Feedback logged for run_id: {dummy_run_id}")

# Another example with negative feedback
dummy_run_negative = client.create_run(
    project_name=os.environ["LANGCHAIN_PROJECT"],
    run_type="llm",
    name="dummy_llm_call_for_feedback_negative",
    inputs={"prompt": "Who won the 2022 World Cup?"},
    outputs={"completion": "Brazil."}, # Incorrect
    status=RunStatus.COMPLETED,
)
dummy_run_negative_id = dummy_run_negative.id

client.create_feedback(
    run_id=dummy_run_negative_id,
    key="user_score",
    score=0.0,
    comment="Incorrect answer, Argentina won.",
)
print(f"Negative feedback logged for run_id: {dummy_run_negative_id}")

このフィードバックは、トレースのフィルタリング、一般的な失敗モードの特定、改善の優先順位付けに使用できます。さらに重要なのは、肯定的なフィードバックを自動的にグラウンドトゥルースの例を生成するために使用できることです。

4.2 グラウンドトゥルースデータセットの作成

LangSmith では、既存の実行からデータセットを作成したり、CSV/JSON をアップロードしたり、プログラムで作成したりできます。

4.2.1 既存の実行から(フィードバック駆動型)

LangSmith でフィードバックスコアによって実行をフィルタリングし、それらをデータセットにエクスポートできます。たとえば、user_score:1.0 の実行をフィルタリングし、UI から「データセットを作成」を選択します。これは、本番環境での使用から評価データセットをブートストラップする強力な方法です。

4.2.2 プログラムによるデータセット作成

新しい機能や特定のテストケースの場合、データセットをプログラムで作成します。

from langsmith import Client
from langsmith.schemas import Example

client = Client()

dataset_name = "RAG Question Answering Evaluation"
dataset_description = "Questions and answers for evaluating the RAG chain."

# Check if dataset exists, create if not
try:
    dataset = client.read_dataset(dataset_name=dataset_name)
    print(f"Dataset '{dataset_name}' already exists.")
except Exception:
    dataset = client.create_dataset(
        dataset_name=dataset_name,
        description=dataset_description,
        data_type="kv", # Key-value pairs
    )
    print(f"Dataset '{dataset_name}' created.")

# Define examples
examples = [
    {"input": "What is the capital of France?", "output": "Paris."},
    {"input": "Who painted the Mona Lisa?", "output": "Leonardo da Vinci."},
    {"input": "What is the highest mountain in the world?", "output": "Mount Everest."},
    {"input": "What is the largest ocean on Earth?", "output": "Pacific Ocean."},
    {"input": "What is the chemical symbol for water?", "output": "H2O."},
]

# Add examples to the dataset
for i, ex in enumerate(examples):
    # Check if example already exists to prevent duplicates on re-run
    # In a real scenario, you might have a more robust check or clear existing examples.
    try:
        client.create_example(
            dataset_id=dataset.id,
            inputs={"input": ex["input"]},
            outputs={"answer": ex["output"]}, # Match the output key of your chain
            metadata={"example_id": f"ex_{i}"}
        )
        print(f"Added example: {ex['input']}")
    except Exception as e:
        if "already exists" in str(e): # Basic check for existing example
            print(f"Example '{ex['input']}' already exists in dataset.")
        else:
            raise e

print(f"Dataset '{dataset_name}' populated with {len(examples)} examples.")

データセットに関する重要な考慮事項:

  • 代表性: データセットが一般的なユーザーのクエリ、エッジケース、既知の失敗モードをカバーしていることを確認します。
  • 多様性: さまざまな質問タイプ(事実、推論、会話)を含めます。
  • グラウンドトゥルースの品質: 例の output(または reference_output)は、間違いなく正しいものでなければなりません。
  • 入力/出力キー: inputs および outputs ディクショナリのキーは、評価するチェーンの期待される入力キーと出力キー(例:RAG チェーンの場合は input と answer)と一致する必要があります。
Advertisement

5. LangSmith を使用した自動評価

自動評価は EDD の基礎です。LangSmith は、単純な完全一致から洗練された LLM-as-a-judge まで、さまざまな評価器を提供します。

5.1 評価の実行

評価を実行するには、以下が必要です。

  1. データセット(セクション 4.2 で作成)。
  2. 「実行可能」(LLM チェーンまたはエージェント)。
  3. 1 つ以上の評価器。
from langsmith import Client
from langsmith.evaluation import evaluate
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
from langchain.chains import create_retrieval_chain
from langchain.chains.combine_documents import create_stuff_documents_chain
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings
from langchain_core.documents import Document
import os

client = Client()

# Re-create the RAG chain for evaluation
llm_eval = ChatOpenAI(model="gpt-4o", temperature=0)
embeddings_eval = OpenAIEmbeddings()

docs_eval = [
    Document(page_content="The capital of France is Paris.", metadata={"source": "wiki"}),
    Document(page_content="Eiffel Tower is in Paris.", metadata={"source": "travel_guide"}),
    Document(page_content="The official language of France is French.", metadata={"source": "wiki"}),
    Document(page_content="Mount Everest is the highest mountain in the world.", metadata={"source": "geography"}),
    Document(page_content="Leonardo da Vinci painted the Mona Lisa.", metadata={"source": "art_history"}),
    Document(page_content="The Pacific Ocean is the largest ocean on Earth.", metadata={"source": "geography"}),
    Document(page_content="Water's chemical symbol is H2O.", metadata={"source": "chemistry"}),
]
vectorstore_eval = FAISS.from_documents(docs_eval, embeddings_eval)
retriever_eval = vectorstore_eval.as_retriever()

system_prompt_eval = (
    "You are an assistant for question-answering tasks. "
    "Use the following pieces of retrieved context to answer the question. "
    "If you don't know the answer, just say that you don't know. "
    "Keep the answer concise."
    "\n\n{context}"
)
prompt_eval = ChatPromptTemplate.from_messages(
    [
        ("system", system_prompt_eval),
        ("human", "{input}"),
    ]
)
question_answer_chain_eval = create_stuff_documents_chain(llm_eval, prompt_eval)
rag_chain_to_evaluate = create_retrieval_chain(retriever_eval, question_answer_chain_eval)

# Define evaluators
# 1. Exact Match Evaluator
# Checks if the predicted output exactly matches the reference output.
# Case-insensitive and whitespace-insensitive by default.
from langsmith.evaluation import ExactMatchEvaluator

exact_match_evaluator = ExactMatchEvaluator(
    criteria={"accuracy": "The predicted output should exactly match the reference output."}
)

# 2. LLM-as-a-Judge Evaluator
# Uses an LLM to assess the quality of the response based on custom criteria.
# Requires an LLM (e.g., OpenAI's gpt-4o).
from langsmith.evaluation import LangChainStringEvaluator

llm_judge_evaluator = LangChainStringEvaluator(
    eval_llm=llm_eval, # Use the same LLM or a more powerful one for evaluation
    config={
        "criteria": {
            "correctness": "Is the submission factually correct and directly answers the question?",
            "conciseness": "Is the submission concise and to the point, avoiding unnecessary verbosity?",
            "relevance": "Is the submission relevant to the question asked?",
        },
        "eval_chain_name": "llm_judge_rag_evaluator",
    },
    # You can also provide a custom prompt for the LLM judge
    # prompt="You are an expert evaluator. Assess the following submission..."
)

# Run the evaluation
print("Starting evaluation run...")
evaluation_results = evaluate(
    llm_or_chain=rag_chain_to_evaluate,
    data=dataset_name, # Name of the dataset created earlier
    evaluators=[
        exact_match_evaluator,
        llm_judge_evaluator,
        # You can add more evaluators here, e.g., "qa", "cot_qa", "embedding_distance"
    ],
    experiment_prefix="rag-chain-v1", # Name for this evaluation run
    metadata={"model_version": "gpt-4o", "retriever_config": "faiss_default"},
    max_concurrency=5, # Adjust based on API rate limits
)

print("\nEvaluation complete. Results available in LangSmith.")
# The evaluation_results object contains a summary and links to the full results in LangSmith.
# You can access metrics like:
# print(evaluation_results["results"][0].evaluator_info)

評価が完了したら、LangSmith プロジェクトの「Testing」セクションに移動します。rag-chain-v1 実験が表示されます。それをクリックして、詳細な結果を表示します。これには以下が含まれます。

  • 各評価器の集計スコア。
  • 入力、参照出力、予測出力、評価器のフィードバックを示す個々の例の結果。
  • 失敗の詳細な調査を可能にする、評価された各実行のトレース。

5.2 評価器の種類とトレードオフ

LangSmith は、さまざまな組み込み評価器を提供しています。

| 評価器の種類 | 説明

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement