Probabilities are generated by a production grade model after 100+ iterations, it is a weighted ensemble of various transformer and tree models (TabDPT + TabPFN3 + XGBoost) using 8000+ data points and curating 31 features such as striking/grappling differentials, Elo, physical stats. This is an ongoing project/research which will be improved with each event as I get to know its shortcomings and strengths.

R38d Performance

Val Log Loss
0.6015
453 fights · selection
Test Log Loss
0.6237
453 fights · holdout
ROC AUC
0.7058
test
Val Accuracy
67.8%
453 fights
Test Accuracy
66.9%
453 fights
Brier Score
0.2160
test
Features
31
Training Set
2,100
fights
Selected Model
WeightedEnsemble_L2

Accuracy by Year (Test)

Period Fights Accuracy
2024 (Dec) 21 76.2%
2025 285 68.4%
2026 (Jan–Jun) 147 62.6%
Full test 453 66.9%

Calibration Curve

Calibration curve

R38d test set: predicted probability vs actual win rate. Closer to the diagonal = better calibrated.

Feature Importance

Feature importance

R38d permutation importance on the test set. Higher = more predictive signal.