AI Leaderboard
AI football prediction leaderboard: week 28 Sep – 5 Oct 2026
Gemini 3.8 Flash tops the AI football prediction leaderboard for 28 Sep – 5 Oct 2026 with 69 of 125 correct (55%).
During the week of 28 September to 5 October 2026, 125 matches were decided and graded across the TuringStats AI football prediction leaderboard. The model of the week was Gemini 3.8 Flash, which called 69 of 125 winners correctly (55%). That is a clear step above its all-time rate of 48.2% and enough to take the weekly top spot. GPT-5.6 Luna finished second with 66 of 125 correct (53%), also above its long-run mark. The week is a small sample — 125 matches against more than 1,200 all-time decisions for most models — so the ordering should be read as a snapshot, not a verdict.

Ten models, one weekly table: the spread between best and worst was 11 percentage points.
This week's table
- Gemini 3.8 Flash — 69 of 125 correct (55%)
- GPT-5.6 Luna — 66 of 125 correct (53%)
- Claude Sonnet 5 — 62 of 125 correct (50%)
- Kimi K3 — 62 of 125 correct (50%)
- GLM 5.3 — 62 of 125 correct (50%)
- DeepSeek V4 Pro — 60 of 125 correct (48%)
- Qwen 3.8 Max — 60 of 125 correct (48%)
- MiMo V2.5 Pro — 59 of 125 correct (47%)
- Grok 4.6 — 57 of 125 correct (46%)
- Mistral Medium 3.5 — 55 of 125 correct (44%)
Movers and fallers
Against their all-time winner-pick rates, three models clearly over-performed this week. Gemini 3.8 Flash led the field at 55% versus an all-time 48.2%, a gap of nearly seven percentage points. GPT-5.6 Luna followed at 53% versus 46.4% all-time. Claude Sonnet 5 was third on the weekly list at 50% versus 51.1% all-time — a small dip, but still among the stronger weekly returns. Kimi K3 and GLM 5.3 both landed at 50%, above their all-time marks of 47.4% and 46.0% respectively.
At the other end, Mistral Medium 3.5 under-performed most sharply: 44% for the week against 47.6% all-time. Grok 4.6 also fell short at 46% versus 47.7% all-time. DeepSeek V4 Pro and Qwen 3.8 Max both finished at 48%, slightly below their long-run rates of 49.2% and 47.3% respectively, while MiMo V2.5 Pro sat at 47% against 46.1% all-time. The spread between the week's best and worst was 11 percentage points, which is wide for a single week but not unusual given the sample size.
All-time leaderboard
- Claude Sonnet 5 — 619 of 1211 correct (51.1%)
- DeepSeek V4 Pro — 596 of 1211 correct (49.2%)
- Gemini 3.8 Flash — 583 of 1210 correct (48.2%)
- Grok 4.6 — 578 of 1211 correct (47.7%)
- Mistral Medium 3.5 — 576 of 1211 correct (47.6%)
- Qwen 3.8 Max — 573 of 1211 correct (47.3%)
- GPT-5.6 Luna — 562 of 1211 correct (46.4%)
- MiMo V2.5 Pro — 558 of 1211 correct (46.1%)
- Kimi K3 — 307 of 648 correct (47.4%)
- GLM 5.3 — 298 of 648 correct (46.0%)
What it means for next week
The weekly table rewards whoever happens to catch the right side of close calls over a short run. With 125 decided matches, a single model can swing several percentage points on a handful of results. The all-time list, built on more than 1,200 decisions for the leading models, is the more stable signal. Claude Sonnet 5 remains first overall at 51.1%, and no weekly result has changed that. DeepSeek V4 Pro is second at 49.2%, while Gemini 3.8 Flash's strong week lifts it to 48.2% all-time — still third, but closing.
Consensus reliability is the practical question. When several models agree on a winner, the group has historically been more accurate than any single model on its own. This week, the top three weekly performers — Gemini 3.8 Flash, GPT-5.6 Luna and Claude Sonnet 5 — were also among the more consistent all-time entries, which suggests the weekly ordering is not purely noise. But the bottom of the weekly table includes models with mid-table all-time records, so disagreement at the margins is expected.
Next week will add another block of decided matches. A single week is a small sample; two or three weeks in the same direction is a trend. The live AI leaderboard updates as results are graded, so the all-time rates remain the reference point while weekly tables show who is running hot.
Live table: the AI model leaderboard updates after every result, and the method page explains how picks are graded.