•22 min read

LangGraph対DSPy:AIエージェントのための宣言型プロンプト最適化とグラフステートマシン

LangGraph対DSPy:AIエージェントのための宣言型プロンプト最適化とグラフステートマシン

2026年に堅牢でプロダクションレベルのAIエージェントを構築するには、オーケストレーションフレームワークとプロンプトエンジニアリングのパラダイムを明確に理解する必要があります。LangGraphとDSPyは、この課題に対する2つの異なる、しかし補完的なアプローチを提示します。LangGraphは、グラフベースの状態機械を介した明示的な状態管理と周期的実行に焦点を当て、DSPyは宣言的なプロンプト最適化とメトリック駆動型コンパイルを推進します。このガイドでは、両者を分析し、それぞれの強み、そしてエンタープライズAI向けの強力なハイブリッドアーキテクチャに統合する方法を実証します。

Audio Briefing
0:00 / 0:00

アーキテクチャパラダイム:LangGraph vs DSPy

LangGraph:明示的な状態機械と周期的実行

LangGraphはLangChainの拡張機能であり、LLM(大規模言語モデル)を使用したステートフルなマルチアクターアプリケーションを構築するためのフレームワークを提供します。そのコアとなる抽象化は、有向非巡回グラフ(DAG)または、より強力な、サイクルを許容する状態グラフです。これにより、複雑で反復的なワークフロー、エージェントループ、およびヒューマン・イン・ザ・ループ(human-in-the-loop)による介入が可能になります。

主な特徴:

  • ノードとエッジ: ワークフローは、エッジで接続された一連のノード(関数、LLM呼び出し、ツール呼び出し)として定義されます。
  • グラフ状態: ノード間で渡される共有の可変状態オブジェクトで、永続的なコンテキストと意思決定を可能にします。
  • 条件付きエッジ: 現在のグラフ状態に基づいて実行を動的にルーティングするロジックです。
  • 周期的実行: ノードを再訪する機能で、エージェントの推論、自己修正、反復的な洗練に不可欠です。
  • ヒューマン・イン・ザ・ループ: 実行を一時停止し、人間の入力を求める明示的なメカニズムです。

LangGraphは、エージェントの動作が複雑で動的なルーティング、反復的な洗練、または純粋に宣言的に表現するのが難しい明示的な状態遷移を必要とする場合に優れています。

DSPy:宣言的なプロンプト最適化とメトリック駆動型コンパイル

DSPy(Declarative Self-improving Language Programs)は、手動のプロンプトエンジニアリングから、プログラムによるメトリック駆動型最適化へとパラダイムを転換します。開発者はプロンプトを手作業で作成する代わりに、LLM呼び出しのシグネチャ(入力フィールド、出力フィールド)を定義し、DSPyは定義されたメトリックに基づいて、これらを特定のLLM向けに最適化されたプロンプトと重みにコンパイルします。

主な特徴:

  • シグネチャ: LLMの入力と出力の抽象的な定義。例:Question -> Answer。
  • モジュール: LLM呼び出しとそのシグネチャをカプセル化する再利用可能なコンポーネント(例:dspy.Chain、dspy.Predict、dspy.Retrieve)。
  • オプティマイザ(コンパイラ): メトリックに対するパフォーマンスを向上させるために、few-shotの例を生成したり、デモンストレーションを再ランク付けしたり、モデルをファインチューニングしたりするアルゴリズム(例:BootstrapFewShot、BayesianSignatureOptimizer)。
  • メトリック: LLM出力の品質を評価するためのユーザー定義関数で、最適化プロセスを推進します。
  • 宣言的: LLMに何をさせるかに焦点を当て、どのようにプロンプトを出すかには焦点を当てません。

DSPyは、プロンプトエンジニアリングプロセスを自動化し、異なるLLMやタスクに適応することで、高品質で堅牢なLLM出力を実現するのに強力です。

アーキテクチャの比較

機能LangGraphDSPyハイブリッド (LangGraph + DSPy)
コア抽象化状態グラフ、ノード、エッジシグネチャ、モジュール、オプティマイザグラフ状態、最適化されたモジュール
制御フロー明示的、命令型状態機械暗黙的、宣言型コンパイル明示的なグラフ、宣言型モジュール
プロンプトエンジニアリング手動、ノード内でのアドホックな処理自動化、メトリック駆動型コンパイル自動化、グラフに統合
状態管理明示的なGraphStateオブジェクトモジュール呼び出し内の暗黙的な処理モジュール状態を持つ明示的なGraphState
周期的ロジックネイティブ、ファーストクラスのサポート直接サポートなしLangGraphを介したネイティブサポート
ヒューマン・イン・ザ・ループネイティブ、明示的なチェックポイント直接サポートなしLangGraphを介したネイティブサポート
最適化手動での反復、デバッグ自動化、メトリック駆動型グラフ内での自動モジュール最適化
最適な用途複雑なマルチエージェントワークフロー、反復推論、動的ルーティング高品質で堅牢な単一ターンまたは連鎖的なLLM呼び出し、プロンプト最適化最適化された堅牢なLLMインタラクションを必要とする複雑なマルチエージェントワークフロー
Advertisement

ハイブリッドアーキテクチャ:DSPyによるコンパイル、LangGraphによるオーケストレーション

これらのフレームワークを組み合わせることで、真の力が発揮されます。DSPyを活用して、個々のLLM呼び出しや複雑な推論ステップに最適なプロンプトシグネチャをコンパイルし、これらの最適化されたDSPyモジュールをLangGraphの状態機械内に組み込むことができます。これにより、LangGraphは全体のエージェントフロー、状態遷移、ツール使用、および人間による介入を管理し、DSPyは各LLMインタラクションの品質と堅牢性を保証します。

例:人間によるレビューを伴うリサーチエージェント

次のようなリサーチエージェントを考えてみましょう。

  1. ユーザーのクエリを受け取る。
  2. 初期検索と要約を実行する。
  3. 潜在的なギャップや曖昧さを特定する。
  4. 必要に応じて、ユーザーに明確化を求める。
  5. 検索を洗練し、最終レポートを生成する。

ここでは、DSPyが検索クエリの生成、要約、ギャップの特定ステップを最適化でき、LangGraphがシーケンスをオーケストレーションし、ヒューマン・イン・ザ・ループによる明確化を処理し、全体の状態を管理します。

ステップ1:コアLLMタスク用のDSPyモジュールを定義する

まず、コアLLM操作用のDSPyシグネチャとモジュールを定義します。

import dspy
from dspy.teleprompt import BootstrapFewShot
from typing import List, Dict, Any

# Configure DSPy with a local LLM (e.g., Ollama) or OpenAI
# For Ollama:
# llm = dspy.Ollama(model="llama3", max_tokens=2000)
# For OpenAI:
llm = dspy.OpenAI(model="gpt-4o-mini", max_tokens=2000)
dspy.settings.configure(lm=llm)

# Define a signature for generating search queries
class GenerateSearchQueries(dspy.Signature):
    """Generate a list of search queries for a given research question."""
    research_question: str = dspy.InputField(desc="The user's research question")
    search_queries: List[str] = dspy.OutputField(desc="A list of relevant search queries")

# Define a signature for summarizing search results
class SummarizeSearchResults(dspy.Signature):
    """Summarize a collection of search results into a concise overview."""
    search_results: List[str] = dspy.InputField(desc="A list of search result snippets")
    summary: str = dspy.OutputField(desc="A concise summary of the search results")

# Define a signature for identifying ambiguities or gaps
class IdentifyGaps(dspy.Signature):
    """Analyze a research summary and identify any ambiguities, missing information, or areas requiring clarification."""
    research_summary: str = dspy.InputField(desc="The current research summary")
    gaps_identified: str = dspy.OutputField(desc="A description of identified gaps or ambiguities, or 'None' if clear")
    requires_clarification: bool = dspy.OutputField(desc="True if user clarification is needed, False otherwise")

# Define DSPy Modules
class SearchQueryGenerator(dspy.Module):
    def __init__(self):
        super().__init__()
        self.generate_queries = dspy.Predict(GenerateSearchQueries)

    def forward(self, research_question: str) -> List[str]:
        prediction = self.generate_queries(research_question=research_question)
        return prediction.search_queries

class SearchResultSummarizer(dspy.Module):
    def __init__(self):
        super().__init__()
        self.summarize = dspy.Predict(SummarizeSearchResults)

    def forward(self, search_results: List[str]) -> str:
        prediction = self.summarize(search_results=search_results)
        return prediction.summary

class GapIdentifier(dspy.Module):
    def __init__(self):
        super().__init__()
        self.identify = dspy.Predict(IdentifyGaps)

    def forward(self, research_summary: str) -> Dict[str, Any]:
        prediction = self.identify(research_summary=research_summary)
        return {
            "gaps_identified": prediction.gaps_identified,
            "requires_clarification": prediction.requires_clarification
        }

# --- Optimization with BootstrapFewShot (example) ---
# In a real scenario, you'd have a dataset of (input, output) pairs
# For demonstration, we'll create a dummy dataset and compile.

# Dummy training data for GenerateSearchQueries
train_data_queries = [
    dspy.Example(research_question="Impact of AI on healthcare", search_queries=["AI in healthcare", "healthcare automation", "AI medical diagnostics"]),
    dspy.Example(research_question="Future of quantum computing", search_queries=["quantum computing trends", "quantum algorithms", "quantum hardware development"]),
]

# Dummy training data for SummarizeSearchResults
train_data_summaries = [
    dspy.Example(search_results=["Snippet 1 about AI", "Snippet 2 about healthcare"], summary="AI is transforming healthcare."),
    dspy.Example(search_results=["Snippet 1 about quantum", "Snippet 2 about future"], summary="Quantum computing holds future promise."),
]

# Dummy training data for IdentifyGaps
train_data_gaps = [
    dspy.Example(research_summary="AI is used in diagnostics.", gaps_identified="Does not specify types of AI or specific diagnostic applications.", requires_clarification=True),
    dspy.Example(research_summary="Quantum computing is a new field.", gaps_identified="None", requires_clarification=False),
]

# Define a simple metric for evaluation (e.g., checking if output is not empty)
def simple_metric(pred, gold, trace=None):
    return bool(pred.search_queries) if 'search_queries' in pred else bool(pred.summary) if 'summary' in pred else bool(pred.gaps_identified)

# Compile the modules
print("Compiling SearchQueryGenerator...")
teleprompter_queries = BootstrapFewShot(metric=simple_metric)
compiled_query_generator = teleprompter_queries.compile(SearchQueryGenerator(), trainset=train_data_queries)
print("SearchQueryGenerator compiled.")

print("Compiling SearchResultSummarizer...")
teleprompter_summaries = BootstrapFewShot(metric=simple_metric)
compiled_summarizer = teleprompter_summaries.compile(SearchResultSummarizer(), trainset=train_data_summaries)
print("SearchResultSummarizer compiled.")

print("Compiling GapIdentifier...")
teleprompter_gaps = BootstrapFewShot(metric=simple_metric)
compiled_gap_identifier = teleprompter_gaps.compile(GapIdentifier(), trainset=train_data_gaps)
print("GapIdentifier compiled.")

# Now, these compiled modules can be used within LangGraph
# For demonstration, let's test them:
# print("\nTesting compiled modules:")
# print(f"Queries: {compiled_query_generator.forward(research_question='Impact of 5G on IoT')}")
# print(f"Summary: {compiled_summarizer.forward(search_results=['5G enables faster IoT', 'IoT devices benefit from low latency'])}")
# print(f"Gaps: {compiled_gap_identifier.forward(research_summary='5G is fast.')}")

ステップ2:LangGraphの状態とノードを定義する

次に、GraphStateと、コンパイルされたDSPyモジュールを使用するノードを定義します。

from typing import TypedDict, Annotated, List, Dict
import operator
from langgraph.graph import StateGraph, END
from langchain_community.tools import DuckDuckGoSearchRun # Example tool

# Define the state for our graph
class ResearchState(TypedDict):
    research_question: str
    search_queries: Annotated[List[str], operator.add]
    search_results: Annotated[List[str], operator.add]
    summary: str
    gaps_identified: str
    requires_clarification: bool
    user_clarification: str
    final_report: str
    iterations: int

# Initialize tools
search_tool = DuckDuckGoSearchRun()

# Define LangGraph nodes
def generate_queries_node(state: ResearchState) -> ResearchState:
    print("---GENERATING QUERIES---")
    question = state["research_question"]
    # Use the compiled DSPy module
    queries = compiled_query_generator.forward(research_question=question)
    return {"search_queries": queries, "iterations": state.get("iterations", 0) + 1}

def perform_search_node(state: ResearchState) -> ResearchState:
    print("---PERFORMING SEARCH---")
    queries = state["search_queries"]
    results = []
    for query in queries:
        print(f"Searching for: {query}")
        # In a real scenario, you'd handle rate limits, errors, etc.
        try:
            result = search_tool.run(query)
            results.append(result)
        except Exception as e:
            print(f"Search failed for '{query}': {e}")
    return {"search_results": results}

def summarize_results_node(state: ResearchState) -> ResearchState:
    print("---SUMMARIZING RESULTS---")
    results = state["search_results"]
    # Use the compiled DSPy module
    summary = compiled_summarizer.forward(search_results=results)
    return {"summary": summary}

def identify_gaps_node(state: ResearchState) -> ResearchState:
    print("---IDENTIFYING GAPS---")
    summary = state["summary"]
    # Use the compiled DSPy module
    gap_info = compiled_gap_identifier.forward(research_summary=summary)
    return {
        "gaps_identified": gap_info["gaps_identified"],
        "requires_clarification": gap_info["requires_clarification"]
    }

def human_in_the_loop_node(state: ResearchState) -> ResearchState:
    print("---HUMAN IN THE LOOP---")
    print(f"Current Summary: {state['summary']}")
    print(f"Gaps Identified: {state['gaps_identified']}")
    user_input = input("Clarification needed. Please provide additional context or guidance (type 'continue' to proceed without further input): ")
    return {"user_clarification": user_input}

def refine_report_node(state: ResearchState) -> ResearchState:
    print("---REFINING REPORT---")
    # This node would typically use another DSPy module for final report generation
    # For simplicity, we'll just combine existing info.
    final_report_content = (
        f"Research Question: {state['research_question']}\n\n"
        f"Summary of Findings:\n{state['summary']}\n\n"
    )
    if state['gaps_identified'] != 'None':
        final_report_content += f"Identified Gaps: {state['gaps_identified']}\n"
    if state['user_clarification']:
        final_report_content += f"User Clarification: {state['user_clarification']}\n"
    final_report_content += "\n--- END OF REPORT ---"
    return {"final_report": final_report_content}

# Define conditional edge logic
def decide_next_step(state: ResearchState) -> str:
    if state["requires_clarification"] and state.get("user_clarification") == None:
        print("---DECISION: CLARIFICATION NEEDED---")
        return "human_review"
    elif state["requires_clarification"] and state.get("user_clarification") != None and state["user_clarification"].lower() != 'continue':
        print("---DECISION: RE-EVALUATE AFTER CLARIFICATION---")
        # If user provided clarification, we might want to re-run search/summarize
        # For this example, we'll just proceed to refine, but in a real system,
        # you'd likely loop back to generate_queries or summarize_results.
        return "refine_report"
    else:
        print("---DECISION: PROCEED TO FINAL REPORT---")
        return "refine_report"

ステップ3:LangGraphワークフローを構築して実行する

最後に、グラフを組み立てて実行します。

# Build the graph
workflow = StateGraph(ResearchState)

workflow.add_node("generate_queries", generate_queries_node)
workflow.add_node("perform_search", perform_search_node)
workflow.add_node("summarize_results", summarize_results_node)
workflow.add_node("identify_gaps", identify_gaps_node)
workflow.add_node("human_review", human_in_the_loop_node)
workflow.add_node("refine_report", refine_report_node)

workflow.set_entry_point("generate_queries")

workflow.add_edge("generate_queries", "perform_search")
workflow.add_edge("perform_search", "summarize_results")
workflow.add_edge("summarize_results", "identify_gaps")

# Conditional edge from identify_gaps
workflow.add_conditional_edges(
    "identify_gaps",
    decide_next_step,
    {
        "human_review": "human_review",
        "refine_report": "refine_report",
    },
)

# After human review, decide if we need to re-evaluate or finalize
workflow.add_conditional_edges(
    "human_review",
    decide_next_step, # Re-use the same decision logic
    {
        "human_review": "human_review", # Loop back if user didn't provide enough info (or if we want to ask again)
        "refine_report": "refine_report",
    },
)

workflow.add_edge("refine_report", END)

app = workflow.compile()

# Run the agent
initial_state = {"research_question": "What are the latest advancements in sustainable energy storage for grid applications?", "search_queries": [], "search_results": [], "summary": "", "gaps_identified": "", "requires_clarification": False, "user_clarification": None, "final_report": "", "iterations": 0}

print("\n--- STARTING RESEARCH AGENT ---")
for s in app.stream(initial_state):
    print(s)
    print("---")

print("\n--- FINAL REPORT ---")
final_state = app.invoke(initial_state)
print(final_state["final_report"])

このハイブリッドアプローチにより、GapIdentifier DSPyモジュールは曖昧さを正確に検出するように最適化され、LangGraphはユーザーに明確化を求め、場合によってはループバックするという複雑なフローを処理します。

プロダクションでの落とし穴とトラブルシューティング

  1. DSPyコンパイルデータの不足:

    • 失敗モード: BootstrapFewShotやその他のオプティマイザが、不十分または低品質なトレーニングデータのためにパフォーマンスが低下する。コンパイルされたプロンプトが最適でなく、LLM出力が一貫しない可能性がある。
    • 修正: 高品質で多様なデモンストレーション例に投資する。重要なモジュールについては、より大きな予算とより堅牢なメトリックを持つBootstrapFewShotWithRandomSearchまたはBayesianSignatureOptimizerの使用を検討する。継続的な評価と再コンパイルのパイプラインを実装する。
    • 実世界のヒント: まず、手動で厳選された少数の例から始める。システムが稼働するにつれて、ユーザーフィードバックや専門家のアノテーションを収集し、トレーニングデータセットを拡張する。
  2. LangGraphの状態管理の複雑さ:

    • 失敗モード: GraphStateが過度に複雑になり、デバッグが困難な状態遷移、競合状態(並行環境で注意深く処理されない場合)、または可変状態による予期しない動作が発生する。
    • 修正: GraphStateを可能な限り簡潔に保つ。偶発的な上書きを防ぐために、リストを蓄積するにはAnnotated[List[str], operator.add]を使用する。明確な命名規則を実装する。複雑な状態については、より良い型強制と検証のためにPydanticモデルの使用を検討する。デバッグを容易にするために、各ノードでの状態変更をログに記録する。
  3. LLMレート制限とコスト超過:

    • 失敗モード: 周期的LangGraph実行、特にDSPyがモジュールごとに複数のLLM呼び出しを行う可能性がある場合、すぐにAPIレート制限に達したり、高額なコストが発生したりする。
    • 修正: 指数関数的バックオフを備えた堅牢な再試行メカニズムを実装する。適切な場合は、同一入力に対するLLM応答をキャッシュする。トークン使用量とコストメトリックを監視する。DSPyの場合、開発およびテスト中の重要度の低いステップには、より小型のファインチューニングされたモデルまたはローカルモデル(例:Ollama経由)の使用を検討する。LangGraphのiterationsカウンターは無限ループを防ぐのに役立つ。
  4. ツール統合の問題(LangGraph):

    • 失敗モード: ツール(例:検索、API呼び出し)がサイレントに失敗したり、不正な形式のデータを返したりして、下流のLLMエラーや不正確なエージェント動作につながる。
    • 修正: ツール呼び出しをtry-exceptブロックでラップする。ツールの入力検証を実装する。ツール出力がLLM消費のために一貫した形式であることを確認する。DSPyモジュールにフィードする前に、LangGraphの専用解析ノードを使用して生のツール出力を処理する。
  5. ハイブリッドシステムのデバッグ:

    • 失敗モード: 問題がLangGraphのオーケストレーションに起因するのか、DSPyのプロンプトコンパイルに起因するのかを特定すること。
    • 修正: コンポーネントを分離する。DSPyモジュールをさまざまな入力で独立してテストし、期待される出力を生成することを確認する。LangGraphのstreamメソッドを使用して、各ステップでの状態変更を観察する。LLMの入力/出力と状態変更を含む、DSPyモジュールとLangGraphノードの両方で詳細なロギングを実装する。DSPyのdspy.settings.trace = Trueはプロンプト生成に関する貴重な洞察を提供できる。

よくある質問

  1. LangGraphとDSPyのどちらを選ぶべきですか?

    • 複雑な多段階推論、動的な制御フロー(条件分岐、ループなど)、ターンをまたぐ明示的な状態管理、ヒューマン・イン・ザ・ループによる介入、または特定のシーケンスでの複数の外部ツールとの統合が必要な場合は、LangGraphを選択してください。
    • 個々のLLM呼び出しまたはLLM呼び出しの短いチェーンの品質と堅牢性を最適化し、手動のプロンプトエンジニアリングの労力を削減し、特定のメトリックに対して高いパフォーマンスを達成することが主な関心事である場合は、DSPyを選択してください。
    • エンタープライズグレードのエージェントの場合、LangGraphでオーケストレーションされたワークフロー内でDSPyを使用して堅牢なLLMインタラクションを実現するハイブリッドアプローチがしばしば優れています。
  2. DSPyはLangGraphワークフロー全体をエンドツーエンドで最適化できますか?

    • 直接はできません。DSPyは、モジュールとして定義された個々のLLM呼び出しまたは呼び出しシーケンスのプロンプトと重みを最適化します。LangGraphのグラフ構造や条件ロジックは最適化しません。ただし、LangGraphノード内のLLMインタラクションを最適化することで、DSPyは間接的にワークフロー全体のパフォーマンスを向上させます。
  3. プロダクション環境でコンパイルされたDSPyモジュールのバージョン管理とデプロイはどのように行いますか?

    • コンパイルされたDSPyモジュール(最適化された内部状態を持つPythonオブジェクト)は、他のモデルアーティファクトと同様に扱います。module.save("path/to/module.json")を使用して保存し、実行時にロードします。これをCI/CDパイプラインに統合し、特定のコンパイル済みバージョンがLangGraphアプリケーションとともにデプロイされるようにします。異なるコンパイル済みバージョンに対してA/Bテストを実装します。
  4. 両方のフレームワークを使用した場合のパフォーマンスへの影響はどうですか?

    • 両方のフレームワークにはオーバーヘッドが伴います。LangGraphは状態管理とグラフトラバーサルのオーバーヘッドを追加します。DSPyはコンパイル時(デプロイごとに1回限りのコスト)と、複雑なfew-shotの例をオンザフライで生成する場合に推論時にもオーバーヘッドを追加する可能性があります。しかし、最適化されたLLMインタラクション(再試行の減少、より正確な出力)と明確なオーケストレーションによるパフォーマンス向上は、特に手動プロンプトが不安定になるような複雑なタスクでは、このオーバーヘッドを上回ることがよくあります。
  5. CrewAIやAutoGenのような他のエージェントフレームワークと比較してどうですか?

    • CrewAIとAutoGenは、マルチエージェントコラボレーションに焦点を当てた高レベルのフレームワークであり、多くの場合、基盤となるオーケストレーションを抽象化しています。これらは内部的にLangChain(したがってLangGraph)または他のメカニズムを使用する可能性があります。LangGraphは、そのようなマルチエージェントシステムを構築するための基本的な状態機械を提供し、よりきめ細かい制御を可能にします。DSPyは、LLMインタラクション層に純粋に焦点を当てています。DSPyを使用してCrewAIまたはAutoGenで定義されたエージェント内のLLM呼び出しを最適化したり、LangGraphを使用して、それらの機能に匹敵するがより明示的な制御を備えたカスタムマルチエージェントシステムを構築したりすることも可能です。
Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement