•6 min read

Practical ML Evaluation: Beyond Accuracy with Precision, Recall, and AUC

Practical ML Evaluation: Beyond Accuracy with Precision, Recall, and AUC

A machine learning model predicting credit card fraud reports 99.9% accuracy. The executive team is thrilled—until they discover that 0.1% of transactions were fraudulent, and the model simply predicted is_fraud = False for every single request. It caught 0% of the fraud while boasting near-perfect accuracy.

This is the Accuracy Paradox. In real-world machine learning—where datasets are imbalanced and the cost of false positives differs wildly from false negatives—raw accuracy is meaningless.

This guide provides a production-grade framework for evaluating classification models, choosing the right metric for your business problem, and tuning decision thresholds in Python.


Audio Briefing
0:00 / 0:00

1. The Confusion Matrix and Core Metrics

                     ┌─────────────────────────────────────────┐
                     │            Actual Reality               │
                     │    Positive (1)   │    Negative (0)     │
┌─────────┬──────────┼───────────────────┼─────────────────────┤
│ Model   │ Pos (1)  │ True Pos (TP)     │ False Pos (FP)      │
│ Predict ├──────────┼───────────────────┼─────────────────────┤
│         │ Neg (0)  │ False Neg (FN)    │ True Neg (TN)       │
└─────────┴──────────┴───────────────────┴─────────────────────┘

From these four values, all primary evaluation metrics derive:

\text{Precision} = \frac{\text{TP}}{\text{TP} + \text{FP}} \quad \text{(When model predicts positive, how often is it right?)}

\text{Recall (Sensitivity)} = \frac{\text{TP}}{\text{TP} + \text{FN}} \quad \text{(Out of all real positives, what percentage did we catch?)}

\text{Specificity} = \frac{\text{TN}}{\text{TN} + \text{FP}} \quad \text{(Out of all real negatives, what percentage did we clear?)}


Advertisement

2. Choosing Between Precision and Recall

Every classification model outputs a continuous probability p \in [0, 1]. Shifting the decision threshold (e.g. from 0.5 to 0.2) increases Recall at the expense of Precision.

Low Threshold (e.g. 0.1) ───────────► High Recall, Low Precision (Catches everything, high noise)
High Threshold (e.g. 0.8) ──────────► High Precision, Low Recall (Only fires on sure things)

When to Prioritize Recall (Minimize False Negatives):

  • Fraud Detection: Missing a $5,000 fraudulent transfer (FN) is far worse than occasionally triggering an SMS verification for a legitimate customer (FP).
  • Disease / Cancer Screening: A missed tumor is fatal; a false positive prompts a harmless follow-up biopsy.
  • Security Threat Detection: Missing an active intrusion is catastrophic.

When to Prioritize Precision (Minimize False Positives):

  • Spam Filters: You would rather see one spam email in your inbox than have your mortgage approval email sent to the Spam folder.
  • Automated Account Bans: Banning an innocent paying customer causes churn and reputation damage.
  • Content Recommendations: Showing irrelevant videos leads users to leave the platform.

3. Combining Metrics: F_1 vs F_\beta Score

The standard F_1 score is the harmonic mean of Precision and Recall, balancing both equally:

F_1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}

When your business prioritizes one over the other, use the F_\beta Score:

F_\beta = (1 + \beta^2) \cdot \frac{\text{Precision} \cdot \text{Recall}}{(\beta^2 \cdot \text{Precision}) + \text{Recall}}

  • \beta = 2.0 (F_2 Score): Weights Recall 2x higher than Precision (ideal for fraud / medical diagnostics).
  • \beta = 0.5 (F_{0.5} Score): Weights Precision 2x higher than Recall (ideal for search & spam filtering).

4. ROC-AUC vs Precision-Recall AUC (PR-AUC)

Both metrics evaluate the model across all possible decision thresholds (from 0.0 to 1.0), but they behave very differently on imbalanced data:

ROC Curve:   Plot of True Positive Rate (Recall) vs False Positive Rate (FPR)
PR Curve:    Plot of Precision vs Recall
from sklearn.metrics import roc_auc_score, average_precision_score

# ROC-AUC is misleadingly high on imbalanced datasets!
roc = roc_auc_score(y_true, y_pred_prob)          # e.g., 0.985 (looks amazing)

# PR-AUC (Average Precision) reflects real-world rare class performance
pr_auc = average_precision_score(y_true, y_pred_prob) # e.g., 0.620 (reveals true difficulty)

Rule of Thumb:

  • If classes are roughly balanced (40/60 to 50/50), use ROC-AUC.
  • If the positive class is rare (< 5% of dataset), PR-AUC (Precision-Recall AUC) is the only trustworthy metric. ROC-AUC will be inflated by the massive count of True Negatives.

Advertisement

5. Probability Calibration: Can You Trust the Numbers?

Many modern classifiers (especially XGBoost, LightGBM, and Deep Neural Networks) output probabilities that are uncalibrated. If a model assigns a probability of 0.80 to 100 users, exactly 80 of them should be positive. If only 40 are positive, the model is overconfident.

Measuring and Fixing Calibration

from sklearn.calibration import CalibratedClassifierCV, calibration_curve
from sklearn.metrics import brier_score_loss
import lightgbm as lgb

# Train raw classifier
base_model = lgb.LGBMClassifier()
base_model.fit(X_train, y_train)

# Calculate Brier Score (lower is better; 0 = perfect calibration)
raw_brier = brier_score_loss(y_test, base_model.predict_proba(X_test)[:, 1])

# Calibrate using Isotonic Regression or Platt Scaling (Sigmoid)
calibrated_model = CalibratedClassifierCV(
    estimator=base_model,
    method='isotonic', # or 'sigmoid' for smaller datasets (< 1000 samples)
    cv='prefit'
)
calibrated_model.fit(X_val, y_val)

calibrated_brier = brier_score_loss(y_test, calibrated_model.predict_proba(X_test)[:, 1])
print(f"Brier score improved from {raw_brier:.4f} to {calibrated_brier:.4f}")

6. Complete Python Implementation: Optimal Threshold Tuning

Here is how to train a classifier on imbalanced data and find the optimal decision threshold based on business costs (Cost(FN) = $100, Cost(FP) = $5):

import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import precision_recall_curve

# Generate synthetic imbalanced dataset (1% positive class)
X, y = make_classification(
    n_samples=50_000, n_features=20, weights=[0.99, 0.01], random_state=42
)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, stratify=y)

# Fit model
clf = HistGradientBoostingClassifier(class_weight='balanced', random_state=42)
clf.fit(X_train, y_train)

# Predict continuous probabilities
y_probs = clf.predict_proba(X_test)[:, 1]

# Calculate Precision-Recall curve
precisions, recalls, thresholds = precision_recall_curve(y_test, y_probs)

# Define Business Cost Function
COST_FALSE_NEGATIVE = 100.0  # Missed fraud cost
COST_FALSE_POSITIVE = 5.0    # User friction cost

total_costs = []
for t in thresholds:
    y_pred = (y_probs >= t).astype(int)
    fn = np.sum((y_test == 1) & (y_pred == 0))
    fp = np.sum((y_test == 0) & (y_pred == 1))
    cost = (fn * COST_FALSE_NEGATIVE) + (fp * COST_FALSE_POSITIVE)
    total_costs.append(cost)

# Find optimal threshold minimizing business loss
best_idx = np.argmin(total_costs)
optimal_threshold = thresholds[best_idx]

print(f"Default 0.5 Threshold Cost:  ${total_costs[np.abs(thresholds - 0.5).argmin()]:,.2f}")
print(f"Optimized Threshold ({optimal_threshold:.3f}) Cost: ${total_costs[best_idx]:,.2f}")

Summary Evaluation Cheatsheet

ScenarioPrimary Metricsecondary Metric
Balanced Binary ClassificationROC-AUCAccuracy / F_1
Fraud / Anomaly Detection (< 2% Pos)PR-AUC (Average Precision)F_2 Score & Cost Curve
Search & Content RetrievalPrecision@K / Mean Average Precision (MAP)NDCG@K
Probability-Sensitive Systems (Bidding / Risk)Brier Score & Expected Calibration Error (ECE)Log-Loss

You Might Also Like

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement