Every card since November 16, 2024: 59 events, 1296 graded picks, both models, nothing removed. 452 of those picks carry a de-vigged sportsbook price (10 cards), which is what the betting section below is built on.
59.8%
Claude, 373 of 624
57.1%
GPT, 384 of 672
61.2%
Both agree, 495 fights
69.8%
Double locks (75%+), 43 fights
Accuracy over time
ClaudeGPTMarket favorite
Running accuracy on the winner after each card. The dashed line is the sportsbook favorite on the same bouts where a price was captured: the bar any pick has to beat. An uninformed coin flip has 50% expected accuracy; realized results fluctuate.
What a confidence number is worth
Stated confidence
Claude actually hit
GPT actually hit
Says 0-59%
48.1% 54
60.0% 15
Says 60-64%
51.8% 110
52.0% 102
Says 65-69%
62.5% 216
55.4% 249
Says 70-74%
60.4% 144
55.1% 178
Says 75-79%
66.7% 81
67.6% 102
Says 80-100%
73.7% 19
65.4% 26
When they disagree, Claude is right
53.7% 123
46.3% 123
A model that says 70 and hits 60 is overconfident. These descriptive hit rates are uncertain, especially in small samples. Trading requires a probability that also holds up against the market on later events.
Main card, prelims, locks
Model
Main card
Prelims
Locks (75%+)
Brier
Claude
64.4% 132/205
57.5% 241/419
68.0% 68/100
0.244
GPT
58.9% 132/224
56.3% 252/448
67.2% 86/128
0.259
Brier is the squared error of the stated probability, lower is better; 0.250 is what guessing 50% on everything scores.
Against the market
Picks
Bets
Hit
Avg price
ROI
Market favoritereference
84
66.7%
67%
-2.9%
Claude, pick was the favorite
53
71.7%
68%
+2.0%
Claude, pick was the underdog
16
43.8%
38%
+8.4%
GPT, pick was the favorite
47
68.1%
69%
-4.0%
GPT, pick was the underdog
15
46.7%
38%
+17.7%
Hypothetical one-contract trades at de-vigged sportsbook probabilities with an assumed 7% × p × (1-p) fee. These are not executable Kalshi prices or returns. Claude: Brier 0.221 against the market's 0.193, best blend weight 0.2; GPT: Brier 0.240 against the market's 0.205, best blend weight 0. A zero selected model weight indicates no benefit from that blend on its training history; it does not establish future performance.
Hypothetical $1,000 bankroll
Claude favorites, conservativeGPT favorites, conservativeEvery Claude pick, conservativeMarket favorite, flat 2%
Card by card since Nov 24, using sportsbook reference prices, configured sizing rules and assumed fees. Calibration uses earlier event dates only. Real spreads, available depth, partial fills and Kalshi execution are absent.
Configurations ranked in hindsight
Bets
Hit
ROI
$1,000 became
All GPT picks · Only when there is a calibrated edge · Aggressive
18
50%
+19.6%
$1,287
All GPT picks · Only when there is a calibrated edge · Medium
18
50%
+21.0%
$1,150
Both models agree · Only when there is a calibrated edge · Aggressive
13
62%
+14.7%
$1,149
All Claude picks · Only when there is a calibrated edge · Aggressive
22
55%
+7.5%
$1,123
Both models agree · Every pick · Aggressive
34
71%
+8.2%
$1,113
Market favorite, flat 2% reference
55
67%
-2.9%
$969
The trading desk Record tab shows chronological research, market benchmarks and scenario details. Settings ranked after observing these returns are not independently verified recommendations. Paper and live execution history are tracked separately.
Best and worst nights
Best
10/10ClaudeUFC Fight Night: Burns vs. MoralesMay 17, 2025 · 100%
10/12ClaudeUFC Fight Night: Bonfim vs. BrownNovember 8, 2025 · 83%
10/12GPTUFC 316: Dvalishvili vs. O'Malley 2June 7, 2025 · 83%
9/11ClaudeUFC 317: Topuria vs. OliveiraJune 28, 2025 · 82%
9/11GPTUFC 317: Topuria vs. OliveiraJune 28, 2025 · 82%
Worst
2/11GPTUFC Fight Night: Oliveira vs. GamrotOctober 11, 2025 · 18%
3/13GPTUFC Fight Night: Sterling vs. ZalalApril 25, 2026 · 23%
3/11GPTUFC Fight Night: Burns vs. MalottApril 18, 2026 · 27%
3/10GPTUFC 321: Aspinall vs. GaneOctober 25, 2025 · 30%
4/12GPTUFC 311: Makhachev vs. Tsarukyan 2January 18, 2025 · 33%
Cards with at least six graded fights for that model.