Fine-Tuning DeepSeek R1 with Unsloth & LoRA: Memory-Efficient Reasoning Models

Table of Contents(15 sections)
DeepSeek-R1-Distill models offer a compelling balance of reasoning capabilities and efficiency. Fine-tuning these models for specific domains or tasks often presents VRAM and training throughput challenges, particularly with larger variants. This guide details a production-grade methodology for fine-tuning DeepSeek-R1-Distill using Unsloth and LoRA/QLoRA, leveraging custom Triton kernels for significant VRAM reduction and accelerated training. We will cover data preparation, model training, and deployment considerations, including GGUF and vLLM export.
DeepSeek-R1-Distill Architecture Overview
DeepSeek-R1-Distill is a series of open-source reasoning models, distilled from larger, more capable DeepSeek-R1 models. They are designed to excel in complex reasoning tasks, often employing a <think> token to delineate an explicit chain-of-thought (CoT) reasoning process before producing a final answer. This explicit CoT structure is crucial for interpretability and debugging, and its preservation during fine-tuning is paramount. The models are typically based on a transformer architecture, utilizing multi-head attention and feed-forward networks.
Unsloth: VRAM Efficiency and Speed
Unsloth is a library designed to accelerate and optimize LoRA/QLoRA fine-tuning for large language models. It achieves substantial VRAM savings and speedups through several key innovations:
- Custom Triton Kernels: Unsloth replaces standard PyTorch attention and optimizer kernels with highly optimized Triton implementations. These kernels are specifically designed for LoRA/QLoRA, reducing memory overhead and improving computational efficiency.
- Gradient Checkpointing Optimization: While PyTorch's gradient checkpointing saves memory, Unsloth further optimizes it by only checkpointing necessary activations, reducing recomputation overhead.
- Quantization-Aware Training: Seamless integration with 4-bit and 8-bit quantization (QLoRA) further slashes VRAM requirements.
These optimizations make it feasible to fine-tune models like DeepSeek-R1-Distill on consumer-grade GPUs or with larger batch sizes on professional hardware.
Data Preparation: Preserving <think> CoT
The DeepSeek-R1-Distill models are trained to use a specific <think> token to structure their reasoning. When preparing fine-tuning data, it is critical to maintain this format. Synthetic data generation or careful annotation is often required.
A typical DeepSeek-R1-Distill conversational turn with CoT looks like this:
User: <prompt>
Assistant: <think>Thought process leading to the answer.</think><answer>Final answer.</answer>
For fine-tuning, we need to format our data into conversational turns, ensuring the <think> and <answer> tags are correctly placed.
Example Data Format
Consider a dataset of mathematical reasoning problems. Each entry should be structured as a list of dictionaries, representing a conversation.
[
{
"messages": [
{
"role": "user",
"content": "What is the sum of the first 10 prime numbers?"
},
{
"role": "assistant",
"content": "<think>The first 10 prime numbers are 2, 3, 5, 7, 11, 13, 17, 19, 23, 29. Summing them: 2+3+5+7+11+13+17+19+23+29 = 129.</think><answer>129</answer>"
}
]
},
{
"messages": [
{
"role": "user",
"content": "If a car travels at 60 mph for 2 hours, how far does it travel?"
},
{
"role": "assistant",
"content": "<think>Distance = Speed × Time. Speed = 60 mph, Time = 2 hours. Distance = 60 * 2 = 120 miles.</think><answer>120 miles</answer>"
}
]
}
]
This JSON structure is compatible with datasets library and Unsloth's data loaders.
Fine-Tuning DeepSeek-R1-Distill with Unsloth
This section provides a complete, runnable example for fine-tuning deepseek-ai/deepseek-r1-3b-base using Unsloth. The principles apply to larger DeepSeek-R1-Distill models as well.
Setup
First, install Unsloth and other necessary libraries.
pip install "unsloth[cu121] @ git+https://github.com/unslothai/unsloth.git"
pip install transformers peft accelerate bitsandbytes trl datasets torch
Training Script
import torch
from unsloth import FastLanguageModel
from trl import SFTTrainer
from transformers import TrainingArguments, AutoTokenizer
from datasets import load_dataset
import os
# 1. Configuration
max_seq_length = 2048 # Max sequence length for DeepSeek-R1-Distill
model_name = "deepseek-ai/deepseek-r1-3b-base" # Or deepseek-ai/deepseek-r1-7b-base
dataset_path = "your_synthetic_reasoning_data.json" # Path to your JSON dataset
# 2. Load Model and Tokenizer with Unsloth
# Unsloth automatically handles 4-bit quantization (QLoRA)
# and loads the model with optimized kernels.
model, tokenizer = FastLanguageModel.from_pretrained(
model_name = model_name,
max_seq_length = max_seq_length,
dtype = None, # None for auto detection (bfloat16 if supported, else float16)
load_in_4bit = True, # Enable QLoRA
)
# 3. Configure LoRA Adapters
# Target all linear layers for optimal performance.
model = FastLanguageModel.get_peft_model(
model,
r = 16, # LoRA rank
target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
lora_alpha = 16,
lora_dropout = 0.05,
bias = "none",
use_gradient_checkpointing = "unsloth", # Use Unsloth's optimized gradient checkpointing
random_state = 3407,
max_seq_length = max_seq_length,
)
# 4. Load and Format Dataset
# The dataset should be a JSON file with the structure described above.
# We use `apply_chat_template` to format messages into the model's expected input format.
# DeepSeek models typically use a specific chat template.
# Ensure the tokenizer has a chat template or define one.
# For DeepSeek, it's often similar to:
# {% for message in messages %}{% if message['role'] == 'user' %}{{ '<|User|>
# ' + message['content'] + '<|EOT|>
# ' }}{% elif message['role'] == 'assistant' %}{{ '<|Bot|>
# ' + message['content'] + '<|EOT|>
# ' }}{% endif %}{% endfor %}{% if add_generation_prompt %}{{ '<|Bot|>' }}{% endif %}
# If your tokenizer doesn't have a default, you might need to set it:
# tokenizer.chat_template = "{% for message in messages %}{% if message['role'] == 'user' %}{{ '<|User|>\n' + message['content'] + '<|EOT|>\n' }}{% elif message['role'] == 'assistant' %}{{ '<|Bot|>\n' + message['content'] + '<|EOT|>\n' }}{% endif %}{% endfor %}{% if add_generation_prompt %}{{ '<|Bot|>' }}{% endif %}"
def formatting_prompts_func(examples):
"""
Formats the dataset examples into the model's expected chat template.
Ensures the <think> and <answer> tokens are preserved within the assistant's response.
"""
texts = []
for i in range(len(examples["messages"])):
# Apply chat template to each conversation
# `add_generation_prompt=False` is crucial for fine-tuning to ensure
# the assistant's response is fully included, not just the prompt for generation.
formatted_text = tokenizer.apply_chat_template(
examples["messages"][i],
tokenize=False,
add_generation_prompt=False
)
texts.append(formatted_text)
return { "text" : texts }
dataset = load_dataset("json", data_files=dataset_path, split="train")
dataset = dataset.map(
formatting_prompts_func,
batched = True,
)
# 5. Configure Training Arguments
trainer = SFTTrainer(
model = model,
tokenizer = tokenizer,
train_dataset = dataset,
dataset_text_field = "text", # Field containing the formatted text
max_seq_length = max_seq_length,
dataset_num_proc = os.cpu_count(), # Use all CPU cores for data processing
packing = False, # Set to True for more efficient packing of short sequences
args = TrainingArguments(
per_device_train_batch_size = 2, # Adjust based on VRAM
gradient_accumulation_steps = 4, # Accumulate gradients over 4 steps
warmup_steps = 5,
num_train_epochs = 3,
learning_rate = 2e-4,
fp16 = not torch.cuda.is_bf16_supported(), # Use fp16 if bfloat16 not supported
bf16 = torch.cuda.is_bf16_supported(), # Use bf16 if supported
logging_steps = 1,
optim = "adamw_8bit", # Use 8-bit AdamW optimizer
weight_decay = 0.01,
lr_scheduler_type = "linear",
seed = 3407,
output_dir = "outputs",
report_to = "none", # Disable reporting to W&B etc. for simplicity
),
)
# 6. Train the Model
trainer.train()
# 7. Save the LoRA Adapters
model.save_pretrained("deepseek_r1_lora_adapters")
tokenizer.save_pretrained("deepseek_r1_lora_adapters")
print("Fine-tuning complete. LoRA adapters saved to deepseek_r1_lora_adapters.")
Explanation of Key Parameters:
max_seq_length: Crucial for memory. DeepSeek-R1-Distill models are often trained with larger contexts. Adjust based on your data and VRAM.load_in_4bit = True: Activates QLoRA, loading the base model in 4-bit precision. This is the primary VRAM saving mechanism.target_modules: Specifies which linear layers in the transformer architecture will have LoRA adapters applied. Targetingq_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_projis standard for comprehensive fine-tuning.use_gradient_checkpointing = "unsloth": Leverages Unsloth's optimized gradient checkpointing.formatting_prompts_func: This function is vital. It takes your structured data and converts it into the exact string format the model expects, including the<think>and<answer>tags. Thetokenizer.apply_chat_templatemethod is the recommended way to do this.per_device_train_batch_sizeandgradient_accumulation_steps: Adjust these to fit your GPU's VRAM. A smallerper_device_train_batch_sizecombined withgradient_accumulation_stepsallows for larger effective batch sizes without increasing peak VRAM.optim = "adamw_8bit": Uses an 8-bit AdamW optimizer, further reducing optimizer state memory.
Deployment: GGUF and vLLM Export
After fine-tuning, you'll have LoRA adapters. For deployment, you typically want to merge these adapters back into the base model weights and then convert them to a suitable format.
Merging LoRA Adapters
import torch
from unsloth import FastLanguageModel
from transformers import AutoTokenizer
model_name = "deepseek-ai/deepseek-r1-3b-base"
lora_adapters_path = "deepseek_r1_lora_adapters"
output_merged_path = "deepseek_r1_merged_model"
# Load the base model and tokenizer
model, tokenizer = FastLanguageModel.from_pretrained(
model_name = model_name,
max_seq_length = 2048, # Must match training max_seq_length
dtype = None, # Use the same dtype as during training or float16/bfloat16
load_in_4bit = False, # Load in full precision for merging
)
# Load the LoRA adapters
model = FastLanguageModel.get_peft_model(
model,
r = 16, # Must match training r
target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
lora_alpha = 16,
lora_dropout = 0.05,
bias = "none",
use_gradient_checkpointing = False, # Not needed for inference/merging
random_state = 3407,
max_seq_length = 2048, # Must match training max_seq_length
)
model.load_adapter(lora_adapters_path)
# Merge LoRA adapters and save the full model
model.save_pretrained_merged(output_merged_path, tokenizer, save_method = "merged_16bit")
print(f"Merged model saved to {output_merged_path}")
GGUF Export for CPU/Edge Deployment
GGUF is a format optimized for CPU inference with llama.cpp and its bindings. It supports quantization (e.g., Q4_K_M, Q5_K_M) for further memory reduction.
import os
from transformers import AutoTokenizer, AutoModelForCausalLM
from huggingface_hub import HfApi, create_repo
from pathlib import Path
merged_model_path = "deepseek_r1_merged_model"
output_gguf_path = "deepseek_r1_merged_model_gguf"
quantization_type = "q4_k_m" # Example: q4_k_m, q5_k_m, q8_0
# Ensure llama.cpp is installed and converted.py is available
# You might need to clone llama.cpp and build it:
# git clone https://github.com/ggerganov/llama.cpp.git
# cd llama.cpp && make
# Load the merged model
model = AutoModelForCausalLM.from_pretrained(
merged_model_path,
torch_dtype=torch.float16, # Or bfloat16
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(merged_model_path)
# Save in a format compatible with llama.cpp's convert.py
# This typically means saving as a standard Hugging Face model
model.save_pretrained(output_gguf_path)
tokenizer.save_pretrained(output_gguf_path)
# Convert to GGUF using llama.cpp's convert.py
# Make sure you have llama.cpp cloned and built, and its `convert.py` is accessible.
# Adjust `llama_cpp_dir` to your llama.cpp installation path.
llama_cpp_dir = "/path/to/llama.cpp" # IMPORTANT: Set this path
convert_script = os.path.join(llama_cpp_dir, "convert.py")
if not os.path.exists(convert_script):
raise FileNotFoundError(f"llama.cpp convert.py not found at {convert_script}. "
"Please clone and build llama.cpp.")
# Step 1: Convert PyTorch weights to ggml format (intermediate step for convert.py)
# This step is often implicitly handled by newer convert.py versions,
# but historically involved a separate script or specific arguments.
# For DeepSeek, ensure convert.py supports its architecture.
print(f"Converting {output_gguf_path} to GGUF...")
os.system(f"python {convert_script} {output_gguf_path} --outfile {output_gguf_path}/model.gguf")
# Step 2: Quantize the GGUF model
quantize_script = os.path.join(llama_cpp_dir, "quantize")
if not os.path.exists(quantize_script):
# For older llama.cpp versions, quantize might be part of convert.py or a separate script
# For newer versions, `quantize` is a compiled binary.
# If not found, try `python {convert_script} --quantize {quantization_type} ...`
raise FileNotFoundError(f"llama.cpp quantize binary not found at {quantize_script}. "
"Please build llama.cpp or check its documentation for quantization.")
os.system(f"{quantize_script} {output_gguf_path}/model.gguf {output_gguf_path}/model-{quantization_type}.gguf {quantization_type}")
print(f"GGUF model ({quantization_type}) saved to {output_gguf_path}/model-{quantization_type}.gguf")
# Optional: Upload to Hugging Face Hub
# api = HfApi()
# repo_id = "your_hf_username/deepseek-r1-3b-reasoning-gguf"
# create_repo(repo_id, repo_type="model", exist_ok=True)
# api.upload_file(
# path_or_fileobj=f"{output_gguf_path}/model-{quantization_type}.gguf",
# path_in_repo=f"deepseek-r1-3b-reasoning-{quantization_type}.gguf",
# repo_id=repo_id,
# )
# print(f"GGUF model uploaded to Hugging Face Hub: {repo_id}")
vLLM Export for High-Throughput GPU Inference
vLLM is an optimized inference engine for LLMs, offering continuous batching and PagedAttention for high throughput. It directly uses Hugging Face transformers models.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
merged_model_path = "deepseek_r1_merged_model"
output_vllm_path = "deepseek_r1_vllm_model"
# Load the merged model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
merged_model_path,
torch_dtype=torch.bfloat16, # Use bfloat16 for vLLM if supported, else float16
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(merged_model_path)
# Save the model in a format vLLM can directly load
# vLLM expects a standard Hugging Face model directory.
model.save_pretrained(output_vllm_path)
tokenizer.save_pretrained(output_vllm_path)
print(f"vLLM-compatible model saved to {output_vllm_path}")
# Example vLLM inference (requires vLLM to be installed)
# from vllm import LLM, SamplingParams
#
# llm = LLM(model=output_vllm_path, dtype=torch.bfloat16)
#
# prompts = [
# "<|User|>\nWhat is the capital of France?<|EOT|>\n<|Bot|>",
# "<|User|>\nIf x = 5 and y = 3, what is x + y?<|EOT|>\n<|Bot|>"
# ]
#
# sampling_params = SamplingParams(temperature=0.7, top_p=0.95, max_tokens=128)
#
# outputs = llm.generate(prompts, sampling_params)
#
# for output in outputs:
# prompt = output.prompt
# generated_text = output.outputs[0].text
# print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
Architecture & Tradeoffs
| Feature | Unsloth LoRA/QLoRA | Full Fine-tuning | GGUF (llama.cpp) | vLLM |
|---|---|---|---|---|
| VRAM Usage (Training) | Very Low (4-bit) | Very High (16/32-bit) | N/A | N/A |
| Training Speed | Fast (Triton kernels) | Standard | N/A | N/A |
| Model Size (Disk) | Adapters only (MBs) | Full model (GBs) | Quantized (GBs) | Full model (GBs) |
| Inference Latency | Moderate (adapter merge) | Low | Low (CPU) | Very Low (GPU) |
| Inference Throughput | Moderate | High | Moderate (CPU) | Very High (GPU) |
| Hardware | Consumer GPUs | High-end GPUs | CPU, Edge devices | High-end GPUs |
| Complexity | Low (Unsloth abstraction) | Moderate | Moderate (conversion) | Low (HF compatible) |
| Use Case | Rapid iteration, resource-constrained training | Max performance, large datasets | Local/offline, CPU-only | High-scale API, GPU-only |
Production Gotchas & Troubleshooting
- VRAM OOM during Training:
- Symptom:
CUDA out of memoryerror. - Fixes:
- Reduce
per_device_train_batch_size. - Increase
gradient_accumulation_stepsto compensate for smaller batch size. - Decrease
max_seq_length. - Ensure
load_in_4bit = Trueandoptim = "adamw_8bit"are set. - Verify
use_gradient_checkpointing = "unsloth". - If using
packing=True, ensure your dataset contains sufficiently long sequences or disable packing if sequences are very short, as packing can sometimes increase peak memory.
- Reduce
- Symptom:
- Incorrect
<think>/<answer>Formatting:- Symptom: Model generates unstructured text, ignores
<think>tags, or produces malformed responses. - Fix: Double-check your
formatting_prompts_funcand the raw JSON data. Ensure the<think>and<answer>tags are exactly as the base model expects and are correctly enclosed within the assistant's response in the training data. Thetokenizer.apply_chat_templateis crucial here.
- Symptom: Model generates unstructured text, ignores
- Slow Training:
- Symptom: Training takes significantly longer than expected.
- Fixes:
- Ensure Unsloth is correctly installed with CUDA support (
unsloth[cu121]). - Verify
torch.cuda.is_available()returnsTrue. - Check GPU utilization with
nvidia-smi. Low utilization might indicate a data loading bottleneck (increasedataset_num_proc). - Ensure
packing=Trueif your sequences are short, as it can improve GPU utilization.
- Ensure Unsloth is correctly installed with CUDA support (
- GGUF Conversion Issues:
- Symptom:
convert.pyfails or produces an unusable GGUF file. - Fixes:
- Ensure
llama.cppis cloned and built correctly. - Verify the
convert.pyscript andquantizebinary paths are correct. - Check
llama.cpp's GitHub for the latest conversion instructions for DeepSeek models, as support for new architectures evolves. - Some models require specific
convert.pyversions or flags.
- Ensure
- Symptom:
- vLLM Inference Errors:
- Symptom: vLLM fails to load the model or generates incorrect output.
- Fixes:
- Ensure the merged model is saved in a standard Hugging Face format that vLLM can recognize.
- Verify
torch_dtypeused for loading in vLLM matches the merged model's dtype (e.g.,bfloat16). - Check vLLM's documentation for specific model compatibility or known issues with DeepSeek.
Frequently Asked Questions
- Can I fine-tune DeepSeek-R1-Distill-7B with Unsloth on a single 24GB GPU (e.g., RTX 3090/4090)?
Yes, with Unsloth's 4-bit QLoRA and optimized gradient checkpointing, fine-tuning DeepSeek-R1-Distill-7B is feasible on a single 24GB GPU. You'll likely need to use a small
per_device_train_batch_size(e.g., 1 or 2) and compensate withgradient_accumulation_steps. - How do I ensure the model continues to use the
<think>token after fine-tuning? The most critical step is to ensure your fine-tuning dataset explicitly includes the<think>and<answer>tokens in the assistant's responses, formatted exactly as the base model expects. Thetokenizer.apply_chat_templatefunction, when used correctly withadd_generation_prompt=False, will help enforce this. During inference, you should prompt the model with the same chat template and expect it to generate the<think>block. - What is the performance impact of 4-bit quantization (QLoRA) on DeepSeek-R1-Distill? While 4-bit quantization introduces a slight precision loss, for many reasoning tasks, the performance degradation is minimal and often outweighed by the significant VRAM savings and ability to fine-tune larger models. DeepSeek models are generally robust to quantization. It's recommended to evaluate the fine-tuned 4-bit model against your specific task metrics.
- Can I use Unsloth for full fine-tuning (not just LoRA)?
Unsloth is primarily designed for LoRA/QLoRA fine-tuning. While it provides optimized kernels, its main benefits in VRAM reduction and speed are tied to the parameter-efficient nature of LoRA. For full fine-tuning, you would typically use standard Hugging Face
transformerstraining, which would require significantly more VRAM. - How do I handle very long reasoning chains that exceed
max_seq_length? If your reasoning chains frequently exceedmax_seq_length, you have a few options:- Increase
max_seq_lengthif VRAM allows. - Truncate overly long reasoning chains in your dataset (though this might impact reasoning quality).
- Consider techniques like "sliding window attention" if the model architecture supports it, or fine-tune on a model with a larger native context window. For DeepSeek-R1-Distill, increasing
max_seq_lengthis the most direct approach.
- Increase
This guide provides a robust framework for fine-tuning DeepSeek-R1-Distill models efficiently. By leveraging Unsloth's optimizations and carefully managing data formatting, you can deploy powerful, specialized reasoning models even with limited hardware resources.
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

DeepSeek-R1 & Distilled Reasoning Models: Local vLLM Deployment, Quantization & Architecture
Comprehensive guide covering deepseek-r1 & distilled reasoning models: local vllm deployment, quantization & architecture with production-grade architecture and code examples.
Read more
SGLang vs vLLM: High-Throughput LLM Inference, RadixAttention & Structured Decoding
Comprehensive guide covering sglang vs vllm: high-throughput llm inference, radixattention & structured decoding with production-grade architecture and code examples.
Read more
LangGraph vs CrewAI in 2026: Multi-Agent Orchestration, State Machines & Cyclic DAGs
Comprehensive guide covering langgraph vs crewai in 2026: multi-agent orchestration, state machines & cyclic dags with production-grade architecture and code examples.
Read more