The Method
The objective is better probabilities and positive returns after costs. Picking more winners does not establish a betting edge. The test is whether a forecast adds information beyond the market, on future cards, at prices we could actually trade.
Start with the market
The baseline is the sportsbook's implied probability after removing its margin from both sides of the same bout. Model agreement and confident analysis are candidate signals; they are not proof that the market is wrong. A sportsbook quote is also not an executable Kalshi price.
Compare every candidate with that baseline on the same eligible bouts using Brier score and log loss, where lower is better. A coin flip scores 0.250 on Brier, but beating a coin flip is insufficient when the market already forecasts the favorites well.
Audit what was known before the fight
Join predictions, odds and settled outcomes by event and fighter pair. Keep a single eligible prediction snapshot, validate both sides of the odds, and report excluded or ambiguous rows. Exclude snapshots that cannot be established as pre-event and bouts without a usable result. A missing value stays missing.
UFCStats, opponent history, ratings, age and layoff can help only if their values were available before the forecast. Reconstructing an old fight with today's career totals leaks its outcome into the features. Date-only timestamps cannot establish the exact information available at a particular betting time; the report must retain that limitation.
Learn small, regularized corrections
The research pipeline compares the unchanged market forecast with shrinkage blends and a regularized logistic residual model. The residual model starts from the market's log odds and learns a limited correction from the available prediction signals. Regularization penalizes large corrections so a small sample has less opportunity to manufacture an edge.
Blending weights and penalty strength are selected using earlier cards inside each training window. The market remains a valid winner of model selection. Language-model confidence is never accepted as a measured win probability just because the explanation sounds persuasive.
Test on the next card, never a random split
Validation moves forward through time. Train on earlier event dates, select settings inside that earlier history, and score the next unseen event date. Keep every bout and prediction from an event together. Leaving one event out while training on later events is not a simulation of forecasting the future.
Report the sample size, validation folds, model and market scores, and event-level uncertainty in their difference. A favorable point estimate with a confidence interval spanning zero does not establish improvement. Ranking many settings on the same history produces a hindsight winner, not an independently verified strategy.
Turn probabilities into executable decisions
For a contract paying $1, the break-even probability is its purchase price plus fees per contract. A trade also needs a margin for uncertainty and execution costs. Fresh quotes, matching outcomes, sufficient liquidity, event timing and available bankroll all matter. A positive model-minus-price number alone is insufficient.
The trading desk shows the estimated debit including fees, spending limits, readiness checks and reasons to abstain. Paper plans help collect evidence; they do not prove that orders would fill. Live automation requires verified research and execution readiness. Manual trading has its own order review and current exchange fee estimate. Actual fees and rounding are determined by Kalshi.
Current research evidence
Updated 2026-09-19T22:40:02.145Z · research_only
No proven trading edge. Use paper tracking while collecting timestamped prices, fills and actual fees.
77 eligible bouts across 9 event dates. 2024-12-07 to 2026-09-19. 1173 source rows excluded by the audit. 1 chronological validation folds.
| CANDIDATE | TEST BOUTS | BRIER | MARKET BRIER | LOG LOSS |
|---|---|---|---|---|
| Sportsbook benchmark | 2 | 0.1002 | 0.1002 | 0.3754 |
| Market + strongly shrunk model signals | 2 | 0.0970 | 0.1002 | 0.3676 |
| Market + model signals | 2 | 0.0919 | 0.1002 | 0.3552 |
| Market + lightly shrunk model signals | 2 | 0.0849 | 0.1002 | 0.3381 |
| Model chosen using earlier events only | 2 | 0.1002 | 0.1002 | 0.3754 |
Live automation readiness
- Operator override: KALSHI_LIVE_AUTOBET is set, so live orders are sent for research, claude, gpt, consensus, double_lock, book without validated execution evidence.
- Historical sportsbook probabilities are not executable Kalshi prices.
- No independently validated, fee-adjusted forward Kalshi execution study has been promoted.
- Fewer than 200 held-out bouts across 20 events.
- Held-out probability improvement over the market is not statistically established.
Data limitations
- Dates without verified pre-event timestamps are excluded; same-day records cannot be certified from an event date alone.
- Research benchmarks require the source quote to be at most 15 minutes old when both the prediction and price were available.
- Candidate scores share a holdout; only the adaptive row evaluates prior-only model selection. Confidence intervals are event-date cluster bootstraps, not profit forecasts.
- Current fighter career totals cannot be used as historical pre-fight features.
- Paper orders assume the quoted quantity fills. Slippage, depth and exchange settlements require forward verification.
Causal fight-history model study
8595 eligible bouts, split chronologically into 5286 training, 1309 validation and 2000 untouched test fights. Model selected on validation: boosting.
| MODEL | TEST BRIER | TEST LOG LOSS | ACCURACY |
|---|---|---|---|
| logistic_C0.1 | 0.2393 | 0.6710 | 58.0% |
| logistic_C1 | 0.2393 | 0.6710 | 58.0% |
| forest | 0.2393 | 0.6711 | 59.0% |
| boosting | 0.2396 | 0.6718 | 58.0% |
| elo | 0.2465 | 0.6860 | 55.0% |
Outcome prediction experiment, not a Kalshi profit backtest. Model selection was frozen before the final holdout.
The market comparison: 66 shared bouts across 7 event dates.
| SAME-SAMPLE FORECAST | BRIER | ACCURACY |
|---|---|---|
| Sportsbook benchmark | 0.2073 | 65.2% |
| logistic_C0.1 | 0.2459 | 57.6% |
| logistic_C1 | 0.2460 | 57.6% |
| forest | 0.2401 | 65.2% |
| boosting | 0.2366 | 65.2% |
| elo | 0.2459 | 53.0% |
Same held-out bouts with independently verified pre-event sportsbook prices. Small sample; no Kalshi fill or profit evidence.
Keep three records separate
Prediction accuracy: graded forecasts, calibration and performance relative to the market. Sportsbook proxy simulations: hypothetical trades using historical sportsbook probabilities and an assumed fee model. Kalshi execution: actual quotes, orders, partial fills, fees and settled returns in the account.
The Record tab labels historical simulations and research separately. Your History tab shows paper and live activity separately. No historical accuracy score or proxy profit establishes routine future profitability.