Practical ML Evaluation: Beyond Accuracy with Precision, Recall, and AUC

Table of Contents
A machine learning model predicting credit card fraud reports 99.9% accuracy. The executive team is thrilled—until they discover that 0.1% of transactions were fraudulent, and the model simply predicted is_fraud = False for every single request. It caught 0% of the fraud while boasting near-perfect accuracy.
This is the Accuracy Paradox. In real-world machine learning—where datasets are imbalanced and the cost of false positives differs wildly from false negatives—raw accuracy is meaningless.
This guide provides a production-grade framework for evaluating classification models, choosing the right metric for your business problem, and tuning decision thresholds in Python.
1. The Confusion Matrix and Core Metrics
┌─────────────────────────────────────────┐
│ Actual Reality │
│ Positive (1) │ Negative (0) │
┌─────────┬──────────┼───────────────────┼─────────────────────┤
│ Model │ Pos (1) │ True Pos (TP) │ False Pos (FP) │
│ Predict ├──────────┼───────────────────┼─────────────────────┤
│ │ Neg (0) │ False Neg (FN) │ True Neg (TN) │
└─────────┴──────────┴───────────────────┴─────────────────────┘
From these four values, all primary evaluation metrics derive:
\text{Precision} = \frac{\text{TP}}{\text{TP} + \text{FP}} \quad \text{(When model predicts positive, how often is it right?)}
\text{Recall (Sensitivity)} = \frac{\text{TP}}{\text{TP} + \text{FN}} \quad \text{(Out of all real positives, what percentage did we catch?)}
\text{Specificity} = \frac{\text{TN}}{\text{TN} + \text{FP}} \quad \text{(Out of all real negatives, what percentage did we clear?)}
2. Choosing Between Precision and Recall
Every classification model outputs a continuous probability p \in [0, 1]. Shifting the decision threshold (e.g. from 0.5 to 0.2) increases Recall at the expense of Precision.
Low Threshold (e.g. 0.1) ───────────► High Recall, Low Precision (Catches everything, high noise)
High Threshold (e.g. 0.8) ──────────► High Precision, Low Recall (Only fires on sure things)
When to Prioritize Recall (Minimize False Negatives):
- Fraud Detection: Missing a $5,000 fraudulent transfer (FN) is far worse than occasionally triggering an SMS verification for a legitimate customer (FP).
- Disease / Cancer Screening: A missed tumor is fatal; a false positive prompts a harmless follow-up biopsy.
- Security Threat Detection: Missing an active intrusion is catastrophic.
When to Prioritize Precision (Minimize False Positives):
- Spam Filters: You would rather see one spam email in your inbox than have your mortgage approval email sent to the Spam folder.
- Automated Account Bans: Banning an innocent paying customer causes churn and reputation damage.
- Content Recommendations: Showing irrelevant videos leads users to leave the platform.
3. Combining Metrics: F_1 vs F_\beta Score
The standard F_1 score is the harmonic mean of Precision and Recall, balancing both equally:
F_1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}
When your business prioritizes one over the other, use the F_\beta Score:
F_\beta = (1 + \beta^2) \cdot \frac{\text{Precision} \cdot \text{Recall}}{(\beta^2 \cdot \text{Precision}) + \text{Recall}}
- \beta = 2.0 (F_2 Score): Weights Recall 2x higher than Precision (ideal for fraud / medical diagnostics).
- \beta = 0.5 (F_{0.5} Score): Weights Precision 2x higher than Recall (ideal for search & spam filtering).
4. ROC-AUC vs Precision-Recall AUC (PR-AUC)
Both metrics evaluate the model across all possible decision thresholds (from 0.0 to 1.0), but they behave very differently on imbalanced data:
ROC Curve: Plot of True Positive Rate (Recall) vs False Positive Rate (FPR)
PR Curve: Plot of Precision vs Recall
from sklearn.metrics import roc_auc_score, average_precision_score
# ROC-AUC is misleadingly high on imbalanced datasets!
roc = roc_auc_score(y_true, y_pred_prob) # e.g., 0.985 (looks amazing)
# PR-AUC (Average Precision) reflects real-world rare class performance
pr_auc = average_precision_score(y_true, y_pred_prob) # e.g., 0.620 (reveals true difficulty)
Rule of Thumb:
- If classes are roughly balanced (40/60 to 50/50), use ROC-AUC.
- If the positive class is rare (< 5% of dataset), PR-AUC (Precision-Recall AUC) is the only trustworthy metric. ROC-AUC will be inflated by the massive count of True Negatives.
5. Probability Calibration: Can You Trust the Numbers?
Many modern classifiers (especially XGBoost, LightGBM, and Deep Neural Networks) output probabilities that are uncalibrated. If a model assigns a probability of 0.80 to 100 users, exactly 80 of them should be positive. If only 40 are positive, the model is overconfident.
Measuring and Fixing Calibration
from sklearn.calibration import CalibratedClassifierCV, calibration_curve
from sklearn.metrics import brier_score_loss
import lightgbm as lgb
# Train raw classifier
base_model = lgb.LGBMClassifier()
base_model.fit(X_train, y_train)
# Calculate Brier Score (lower is better; 0 = perfect calibration)
raw_brier = brier_score_loss(y_test, base_model.predict_proba(X_test)[:, 1])
# Calibrate using Isotonic Regression or Platt Scaling (Sigmoid)
calibrated_model = CalibratedClassifierCV(
estimator=base_model,
method='isotonic', # or 'sigmoid' for smaller datasets (< 1000 samples)
cv='prefit'
)
calibrated_model.fit(X_val, y_val)
calibrated_brier = brier_score_loss(y_test, calibrated_model.predict_proba(X_test)[:, 1])
print(f"Brier score improved from {raw_brier:.4f} to {calibrated_brier:.4f}")
6. Complete Python Implementation: Optimal Threshold Tuning
Here is how to train a classifier on imbalanced data and find the optimal decision threshold based on business costs (Cost(FN) = $100, Cost(FP) = $5):
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import precision_recall_curve
# Generate synthetic imbalanced dataset (1% positive class)
X, y = make_classification(
n_samples=50_000, n_features=20, weights=[0.99, 0.01], random_state=42
)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, stratify=y)
# Fit model
clf = HistGradientBoostingClassifier(class_weight='balanced', random_state=42)
clf.fit(X_train, y_train)
# Predict continuous probabilities
y_probs = clf.predict_proba(X_test)[:, 1]
# Calculate Precision-Recall curve
precisions, recalls, thresholds = precision_recall_curve(y_test, y_probs)
# Define Business Cost Function
COST_FALSE_NEGATIVE = 100.0 # Missed fraud cost
COST_FALSE_POSITIVE = 5.0 # User friction cost
total_costs = []
for t in thresholds:
y_pred = (y_probs >= t).astype(int)
fn = np.sum((y_test == 1) & (y_pred == 0))
fp = np.sum((y_test == 0) & (y_pred == 1))
cost = (fn * COST_FALSE_NEGATIVE) + (fp * COST_FALSE_POSITIVE)
total_costs.append(cost)
# Find optimal threshold minimizing business loss
best_idx = np.argmin(total_costs)
optimal_threshold = thresholds[best_idx]
print(f"Default 0.5 Threshold Cost: ${total_costs[np.abs(thresholds - 0.5).argmin()]:,.2f}")
print(f"Optimized Threshold ({optimal_threshold:.3f}) Cost: ${total_costs[best_idx]:,.2f}")
Summary Evaluation Cheatsheet
| Scenario | Primary Metric | secondary Metric |
|---|---|---|
| Balanced Binary Classification | ROC-AUC | Accuracy / F_1 |
| Fraud / Anomaly Detection (< 2% Pos) | PR-AUC (Average Precision) | F_2 Score & Cost Curve |
| Search & Content Retrieval | Precision@K / Mean Average Precision (MAP) | NDCG@K |
| Probability-Sensitive Systems (Bidding / Risk) | Brier Score & Expected Calibration Error (ECE) | Log-Loss |
You Might Also Like
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

Vector Databases for Production RAG (2026): Pinecone vs Qdrant vs Milvus vs pgvector
An architectural benchmark of Pinecone, Qdrant, Milvus, and pgvector for production RAG pipelines: HNSW vs IVFFlat indexing, single-stage filtered search, p95 latency, and memory footprint.
Read more
Python Libraries I'd Bet On for 2026
I burned too many weekends on the wrong Python libraries so you don't have to. Here's what I actually use to ship things that stay working.
Read more
AI Agent Architecture in Practice: Memory, Tool Use, and Failure Modes
A working guide to building AI agent systems in 2026: ReAct loops, vector memory, tool calling patterns, multi-agent coordination, and how to test agents that fail gracefully.
Read more