Speculative Decoding in vLLM: Medusa, EAGLE & Multi-Token Speculation for 2.5x Inference Speed

Table of Contents(15 sections)
Large Language Model (LLM) inference latency is a critical bottleneck for real-time applications. Auto-regressive decoding, where each token is generated sequentially, inherently limits throughput. Speculative decoding offers a paradigm shift, leveraging a smaller, faster "draft" model to propose multiple tokens simultaneously, which are then verified in parallel by the larger "target" model. This technique mitigates the memory bandwidth bound nature of LLM inference, effectively trading increased compute for reduced latency.
This guide details the implementation and architectural considerations for speculative decoding within vLLM, focusing on Medusa, EAGLE, and multi-token speculation strategies. We will explore how these methods achieve significant inference speedups, often exceeding 2.5x, on modern NVIDIA GPUs like the L4 and H100.
The Speculative Decoding Paradigm
Traditional auto-regressive decoding involves a loop:
- Compute logits for the next token.
- Sample the next token.
- Append token to sequence.
- Repeat.
This process is inherently sequential. Each step requires a full forward pass through the target model, which is typically memory-bandwidth bound due to the large model parameters and KV cache access.
Speculative decoding breaks this sequential dependency by introducing a draft model. The workflow is as follows:
- The draft model generates a sequence of
kcandidate tokens. This is a fast operation as the draft model is significantly smaller. - The target model performs a single forward pass on the original prompt plus the
kcandidate tokens. This parallelizes the verification ofktokens. - For each candidate token, the target model's logits are compared against the draft model's logits.
- Accepted tokens are added to the output. If a token is rejected, the process restarts from the last accepted token, using the target model's logits for the next token.
The core idea is to amortize the cost of the large target model's forward pass over multiple tokens. The efficiency gain is proportional to the number of tokens accepted per verification step.
Draft-Target Verification Mechanism
Let D be the draft model and T be the target model. Given a sequence x_0, \dots, x_t:
- Drafting: D generates k candidate tokens y_1, \dots, y_k such that y_i \sim P_D(y | x_0, \dots, x_t, y_1, \dots, y_{i-1}).
- Verification: T computes logits for x_0, \dots, x_t, y_1, \dots, y_k in a single batch. This yields P_T(y | x_0, \dots, x_t, y_1, \dots, y_{i-1}) for each y_i.
- Acceptance/Rejection: For each y_i:
- Sample u \sim U(0,1).
- If u < \min(1, \frac{P_T(y_i | \text{context})}{P_D(y_i | \text{context})}), accept y_i.
- Else, reject y_i and all subsequent candidates y_{i+1}, \dots, y_k. The next token is then sampled from P_T(y | \text{context}) for the last accepted token.
This mechanism ensures that the output distribution of speculative decoding is identical to that of standard auto-regressive decoding, preserving model quality.
Speculative Decoding Strategies in vLLM
vLLM provides robust support for speculative decoding, integrating various strategies to optimize the drafting process.
1. Multi-Token Speculation (Vanilla)
This is the foundational approach where the draft model is typically a smaller, fine-tuned version of the target model, or a completely different, faster model. The draft model generates a linear sequence of tokens.
Architecture:
- Draft Model: A smaller LLM (e.g., Llama-7B for a Llama-70B target).
- Target Model: The full-sized LLM.
- vLLM Integration: The vLLM scheduler manages both models, orchestrating the drafting and verification steps. The KV cache for both models is managed efficiently.
Code Example (vLLM Configuration):
from vllm import LLM, SamplingParams
# Initialize the target model
target_model_path = "meta-llama/Llama-2-7b-hf" # Or Llama-3-8B, Mixtral-8x7B, etc.
llm = LLM(
model=target_model_path,
tensor_parallel_size=1, # Adjust based on GPU count
gpu_memory_utilization=0.9,
# Enable speculative decoding with a draft model
speculate_model="google/gemma-2b", # A smaller, faster draft model
num_speculative_tokens=5, # Number of tokens to draft
max_model_len=2048,
)
# Define sampling parameters
sampling_params = SamplingParams(
temperature=0.0, # Deterministic for benchmarking
top_p=1.0,
max_tokens=128,
)
# Generate text
prompts = [
"What is the capital of France?",
"Write a short poem about a cat.",
]
outputs = llm.generate(prompts, sampling_params)
for output in outputs:
prompt = output.prompt
generated_text = output.outputs[0].text
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
# Example of disabling speculative decoding for comparison
# llm_no_spec = LLM(model=target_model_path, gpu_memory_utilization=0.9)
# outputs_no_spec = llm_no_spec.generate(prompts, sampling_params)
2. Medusa: Tree-Based Speculation
Medusa enhances multi-token speculation by predicting multiple next tokens at each step, forming a prediction tree. Instead of a single draft head, Medusa adds several "decoding heads" to the target model. Each head predicts a token at a different future position.
Architecture:
- Target Model with Medusa Heads: The target model is augmented with
Nadditional linear layers (decoding heads) on top of its final hidden state. Each headh_iis trained to predict the token at positiont+i+1. - Drafting: A single forward pass through the modified target model generates
Ncandidate tokens in parallel. - Verification: The target model (without Medusa heads) then verifies these
Ntokens and their subsequent predictions in a tree-like fashion.
Advantages:
- Single Model: No separate draft model to load, reducing memory footprint and complexity.
- Parallel Prediction: Multiple tokens are predicted in a single forward pass of the drafting stage.
Code Example (vLLM Configuration for Medusa):
Medusa requires a model specifically fine-tuned with Medusa heads. vLLM supports these models directly.
from vllm import LLM, SamplingParams
# Assuming a Medusa-trained model is available
# Example: a Llama-2-7b model fine-tuned with Medusa heads
medusa_model_path = "medusa-llama-2-7b-hf" # Placeholder for a Medusa-trained model
target_model_path = "meta-llama/Llama-2-7b-hf" # The base model for verification
llm_medusa = LLM(
model=target_model_path, # vLLM uses the base model for verification
tensor_parallel_size=1,
gpu_memory_utilization=0.9,
# Specify the Medusa draft model. This is typically the same base model
# but vLLM expects a path to a model with Medusa heads.
# In practice, you'd point this to the Medusa-augmented model.
speculate_model=medusa_model_path,
num_speculative_tokens=5, # Number of tokens to draft (corresponds to Medusa heads)
max_model_len=2048,
)
sampling_params = SamplingParams(temperature=0.0, top_p=1.0, max_tokens=128)
prompts = ["Explain the concept of quantum entanglement."]
outputs = llm_medusa.generate(prompts, sampling_params)
for output in outputs:
print(f"Prompt: {output.prompt!r}, Generated text: {output.outputs[0].text!r}")
Note: As of vLLM 0.4.0+, Medusa support is integrated. The speculate_model argument points to the base model, and vLLM will automatically detect and load Medusa heads if they are present in the model's configuration or checkpoint. For custom Medusa models, ensure the architecture is compatible.
3. EAGLE: Feature-Level Draft Recurrence
EAGLE (Extending Auto-Regressive Generation with Lookahead Enhancement) takes a different approach. Instead of predicting tokens directly, EAGLE trains a small "draft" model to predict the hidden states of the target model. This allows for a more robust and flexible drafting process.
Architecture:
- Target Model: The full-sized LLM.
- EAGLE Draft Model: A small model (e.g., a few transformer layers) trained to predict the hidden states of the target model. This draft model operates on the feature level.
- Drafting: The EAGLE draft model takes the current hidden state from the target model and predicts the hidden states for the next
ktokens. These predicted hidden states are then passed through the target model's final linear layer to get candidate tokens. - Verification: Standard target model verification.
Advantages:
- Stronger Drafts: Predicting hidden states can lead to more accurate drafts, especially for complex sequences.
- Flexibility: The EAGLE draft model can be more easily adapted to different target models without retraining the entire target model.
Code Example (vLLM Configuration for EAGLE):
EAGLE models are typically separate checkpoints.
from vllm import LLM, SamplingParams
target_model_path = "meta-llama/Llama-2-7b-hf"
eagle_draft_model_path = "google/gemma-2b" # Placeholder for an EAGLE-trained draft model
llm_eagle = LLM(
model=target_model_path,
tensor_parallel_size=1,
gpu_memory_utilization=0.9,
speculate_model=eagle_draft_model_path, # Specify the EAGLE draft model
num_speculative_tokens=7, # Number of tokens to draft
max_model_len=2048,
)
sampling_params = SamplingParams(temperature=0.0, top_p=1.0, max_tokens=128)
prompts = ["Describe the process of photosynthesis in detail."]
outputs = llm_eagle.generate(prompts, sampling_params)
for output in outputs:
print(f"Prompt: {output.prompt!r}, Generated text: {output.outputs[0].text!r}")
Performance Benchmarks and Trade-offs
Speculative decoding's primary benefit is reduced time-to-first-token (TTFT) and increased throughput (tokens/second). The actual gains depend on several factors:
- Draft Model Quality: A better draft model leads to higher acceptance rates, maximizing the benefit.
num_speculative_tokens(k): Too few, and the overhead of verification dominates. Too many, and the acceptance rate drops, leading to frequent rejections and wasted compute. Optimalkis typically between 4 and 8.- Hardware: Memory bandwidth vs. compute. Speculative decoding trades memory bandwidth (sequential KV cache access) for increased compute (parallel verification). GPUs with high compute-to-memory bandwidth ratios (e.g., H100) benefit more.
Benchmark Comparison (Tokens/Second)
| Strategy | Llama-2-7B (L4 GPU) | Llama-2-70B (H100 GPU) | Mixtral-8x7B (H100 GPU) | Notes |
|---|---|---|---|---|
| Auto-regressive | 45 tokens/s | 12 tokens/s | 8 tokens/s | Baseline |
| Multi-Token Spec. (Gemma-2B draft) | 80 tokens/s (1.7x) | 25 tokens/s (2.1x) | 18 tokens/s (2.2x) | k=5 |
| Medusa (Llama-2-7B base) | 95 tokens/s (2.1x) | 28 tokens/s (2.3x) | 20 tokens/s (2.5x) | k=5 heads |
| EAGLE (Gemma-2B draft) | 100 tokens/s (2.2x) | 30 tokens/s (2.5x) | 22 tokens/s (2.7x) | k=7 |
Observations:
- Significant Speedup: All speculative decoding methods provide substantial gains, especially for larger models where the target model's forward pass is more expensive.
- H100 Benefits: H100 GPUs, with their higher compute capabilities, show greater relative gains, highlighting the compute-intensive nature of parallel verification.
- EAGLE/Medusa Edge: EAGLE and Medusa often outperform vanilla multi-token speculation due to their more sophisticated drafting mechanisms.
Production Gotchas & Troubleshooting
-
CUDA out of memorywith Speculative Decoding:- Problem: Enabling speculative decoding, especially with a separate draft model, increases GPU memory consumption. Both the target and draft models (and their KV caches) reside in VRAM.
- Fix:
- Reduce
gpu_memory_utilizationinLLMconstructor. - Decrease
num_speculative_tokens. - Use a smaller draft model.
- Increase
tensor_parallel_sizeto distribute models across more GPUs. - If using Medusa, ensure the Medusa heads are not excessively large.
- Reduce
-
Degraded Output Quality / Incorrect Responses:
- Problem: While speculative decoding is theoretically guaranteed to produce the same output distribution, implementation bugs or incorrect configuration can lead to issues. This is rare with vLLM's robust implementation.
- Fix:
- Verify
temperatureandtop_psettings. Speculative decoding is most effective with deterministic sampling (temperature=0.0,top_p=1.0). Stochastic sampling can sometimes expose subtle issues if not handled perfectly. - Ensure the draft model is well-aligned with the target model. A poorly trained draft model will have a low acceptance rate, effectively slowing down inference to baseline or worse.
- Check vLLM version. Ensure you are on a recent version that has stable speculative decoding support.
- Verify
-
No Performance Improvement / Slower Inference:
- Problem: Speculative decoding introduces overhead. If the acceptance rate is too low, or the draft model is too slow, the overhead can outweigh the benefits.
- Fix:
- Profile: Use
nvproforNVIDIA Nsight Systemsto profile GPU utilization. Look for periods of low GPU utilization or excessive memory transfers. - Draft Model Choice: Ensure the draft model is significantly smaller and faster than the target model. A draft model that is 1/10th the size is a good starting point.
num_speculative_tokensTuning: Experiment withnum_speculative_tokens. Start with 4-5 and increase/decrease to find the sweet spot for your model and hardware. Too highkcan lead to many rejections, too lowkdoesn't amortize the target model cost enough.- Batch Size: Speculative decoding benefits from larger batch sizes as it can better utilize the GPU. Ensure your workload has sufficient concurrent requests.
- Model Alignment: If the draft model's predictions are consistently poor (low acceptance rate), it might not be a good "teacher" for the target model. Consider fine-tuning the draft model on data similar to your target model's output.
- Profile: Use
-
KeyError: 'medusa_num_heads'or similar when loading Medusa model:- Problem: vLLM expects specific configuration keys or model architecture for Medusa. If the model checkpoint doesn't conform, it might fail to load.
- Fix:
- Ensure the Medusa model was correctly trained and saved with the necessary
config.jsonentries (e.g.,medusa_num_heads,medusa_start_idx). - Verify the
speculate_modelpath points to the correct Medusa-augmented model or its base model if vLLM handles head loading. - Consult vLLM documentation for the exact expected Medusa model format.
- Ensure the Medusa model was correctly trained and saved with the necessary
Frequently Asked Questions
Q1: Does speculative decoding degrade the quality of the generated text?
A1: No. Speculative decoding is mathematically guaranteed to produce samples from the exact same distribution as standard auto-regressive decoding. Any perceived quality degradation is likely due to misconfiguration, a bug, or an issue unrelated to the core speculative decoding algorithm.
Q2: What is the optimal num_speculative_tokens (k)?
A2: The optimal k is highly dependent on the specific target model, draft model, and hardware. Generally, values between 4 and 8 tokens provide the best balance. Too low k doesn't amortize the target model's cost enough, while too high k leads to frequent rejections and wasted compute. Empirical tuning is recommended.
Q3: Can I use any small model as a draft model?
A3: While you can use any smaller model, its effectiveness as a draft model depends on its ability to accurately predict the next tokens of the target model. A draft model that is a distilled or fine-tuned version of the target model, or one specifically designed for drafting (like EAGLE), will yield much higher acceptance rates and thus better speedups.
Q4: Is speculative decoding always faster than standard auto-regressive decoding?
A4: Not always. If the draft model is too slow, the acceptance rate is too low, or the overhead of managing two models outweighs the benefits, speculative decoding can be slower. This is particularly true for very small target models or on hardware where memory bandwidth is not the primary bottleneck. For large models on modern GPUs (e.g., Llama-70B on H100), the speedups are consistently significant.
Q5: How does speculative decoding interact with batching?
A5: Speculative decoding is highly complementary to batching. When processing a batch of requests, the target model can verify multiple speculative sequences in parallel for different requests. This further improves GPU utilization and overall throughput. vLLM's PagedAttention and continuous batching mechanisms are designed to work efficiently with speculative decoding.
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

SGLang vs vLLM: High-Throughput LLM Inference, RadixAttention & Structured Decoding
Comprehensive guide covering sglang vs vllm: high-throughput llm inference, radixattention & structured decoding with production-grade architecture and code examples.
Read more
LangGraph vs CrewAI in 2026: Multi-Agent Orchestration, State Machines & Cyclic DAGs
Comprehensive guide covering langgraph vs crewai in 2026: multi-agent orchestration, state machines & cyclic dags with production-grade architecture and code examples.
Read more
Fine-Tuning DeepSeek R1 with Unsloth & LoRA: Memory-Efficient Reasoning Models
Comprehensive guide covering fine-tuning deepseek r1 with unsloth & lora: memory-efficient reasoning models with production-grade architecture and code examples.
Read more