Probabilities are generated by a production grade model after 100+ iterations, it is a weighted ensemble of various frontier foundation transformers and tree models (TabDPT 1.3 + TabPFN-3 + TabICLv2 + CatBoost + LightGBM) using 8,700+ data points and curating 39 mathematical features such as opponent-adjusted decayed performance (AdjPerf), striking/grappling reliability, Elo trend, and physical differentials. This is an ongoing project/research which will be improved with each event as I get to know its shortcomings and strengths.

Aditya Parkhe

R51 Performance (AutoGluon 1.6.1 + TabDPT 1.3)

Val Log Loss
0.5738
406 fights · selection
Test Log Loss
0.6210
416 fights · holdout
ROC AUC
0.7093
test
Val Accuracy
68.0%
406 fights
Test Accuracy
65.9%
416 fights
Brier Score
0.2150
test
Features
39
Training Set
1,914
fights
Selected Model
WeightedEnsemble_L2

Accuracy by Year (Test)

Period Fights Accuracy
2024 (Dec) 8 75.0%
2025 271 66.8%
2026 (Jan–Jun) 137 63.5%
Full test 416 65.9%

Calibration Curve

Calibration curve

R51 test set: predicted probability vs actual win rate. Closer to the diagonal = better calibrated.

Feature Importance

Feature importance

R51 permutation importance on the test set. Higher = more predictive signal.