MLB Predictor
Predicts the winner, the combined score and the run line for every game, before it starts. It places no bets and holds no positions — this page tracks whether it is any good. Hover any ? for an explanation.
Today’s best pick
Log in to see the model’s highest-confidence pick for today, plus the full slate.
Log inSince inception
Model accuracy
Is our model adding anything??
Our published number is part betting-market price, part our own model. Here are the same games scored three ways, so you can see which part is doing the work. If the published number can’t beat the market price on its own, our model isn’t earning the 20% share it is given and that share should come down. This is the single most important row on the page, and also the one that needs the largest sample before it means anything.
| Scored on | Accuracy | vs. published |
|---|---|---|
| What we publish | 0.2332 | — |
| Market price on its own | 0.2339 | +0.0007 |
| Our model on its own | 0.2342 | +0.0010 |
| Always saying 50/50 | 0.2500 | +0.0168 |
Does this model beat the market? No.
We tested this properly on 4,866 completed games across two full seasons, predicting each one using only games that had already finished. Here is everything the inputs can tell you:
| Predicting each game from | Accuracy | Correct |
|---|---|---|
| Nothing — just “home teams win 53%” | 0.2490 | 53.2% |
| How good the two teams have been | 0.2435 | 56.4% |
| Also who is pitching | 0.2471 | 54.6% |
Knowing everything about both teams and both starting pitchers gets you to 56.4%. That is the ceiling for this kind of model, and baseball is simply like that — a single game is close to a coin flip.
The betting market clears that ceiling comfortably, because it also knows the lineup card, the weather, and whoever is quietly unavailable tonight. On our own settled predictions it carried about ten times more information than team records and pitching can supply.
One thing our model does do well: it is honest about its own uncertainty. When it says 70%, that happens about 70% of the time — measurably better calibrated than the market itself. It is not wrong. It just knows less, and no amount of adjusting the settings creates knowledge that is not in the inputs.
So read the picks below as a public experiment, not as an edge. We publish them, we score them, and we will say plainly if that ever changes. The honest way to beat this market is better information — lineups, weather, bullpen fatigue — or reacting faster than it does, not a cleverer weighting of what everyone already has.
If you had followed the picks?
One unit of risk per pick, with Kalshi’s fees taken out. Read the second table before the first — the first uses midpoint prices, and the midpoint is not a price anyone will trade with you at.
| Market | Picks | Units | Per pick |
|---|---|---|---|
| Moneyline | 432 | +6.30u | +0.015u |
| Total | 4513 | -22.64u | -0.005u |
| Run line | 2466 | -38.99u | -0.016u |
| All, at midpoint prices | 7411 | -55.33u |
At prices you could actually have traded
For 89% of these picks we recorded the real order book, so we can price them honestly — you pay the ask to buy, not the midpoint. Same picks, both ways:
| Priced at | Picks | Units | Per pick |
|---|---|---|---|
| The midpoint (flattering) | 6605 | -72.17u | -0.011u |
| What you would really have paid | 6605 | -121.84u | -0.018u |
Treat the honest row as the real one. Fees and the gap between the midpoint and the asking price are not rounding — on this sample they are the whole result. An earlier version of this card ignored both and showed a far better number; it was wrong. And these are only 7411 picks across a few dozen games, with up to eleven strikes priced on a single game, so they are nowhere near 7411 independent bets.
Latest model updates?
What has changed about how these picks are made, and what the evidence was. Changes we considered and decided against are here too.
-
02 Sep 2026Backtested the model on 4,866 games. Verdict: it does not beat the market, and it probably cannot with the information it has.
We had been judging this model on five days of results. So we tested it properly: every completed game of the 2024 and 2025 seasons, predicting each one using only games that had already finished. Knowing how good both teams have been gets you to 56.4% correct. Adding who is pitching does not improve on that. Predicting nothing at all — just “home teams win 53% of the time” — gets 53.2%. That is the ceiling for this kind of model, and baseball is simply like that: one game is close to a coin flip. The betting market clears that ceiling because it also knows the lineup card, the weather, and who is quietly unavailable. The one thing our model does better than the market is know its own uncertainty — when it says 70%, it happens about 70% of the time, and by that measure it is the more honest of the two. But being honest about a weaker opinion is not an edge. We have moved the published number further toward the market price (now 85%, from 80%) and we are not going to keep adjusting settings to disguise this: no weighting creates information that is not in the inputs.
This is the most thorough test we have run — 4,866 games rather than a few hundred — and it is the one we trust. Treat the picks as a public experiment, not as an edge. Beating this market needs better information (lineups, weather, bullpen fatigue) or faster reactions, and we will say so plainly if that ever changes.
-
02 Sep 2026Rebuilt the daily picks: one per game, no moneylines, and no longer ranked by confidence
The old card ranked every priced contract by edge and showed the top twelve. That put the same game on the list five times — one low-scoring opinion produces a positive edge on every under strike at once — so on 1 September the twelve picks shown were really three games, nine of them unders and not a single over. It read as a model that never backs the over. It is not: across 407 settled totals it picked the over 56% of the time. The ranking was manufacturing the bias. Worse, the ranking itself was backwards: taking only our highest-edge picks returned LESS than taking them all, measured two separate ways. The card is now at most seven picks, one per game, main lines only, no moneylines, shown in start-time order rather than ranked.
We do not yet have a way of choosing our best picks that survives testing, so we are not pretending to. Ordering is neutral until one does.
-
01 Sep 2026Corrected the “if you had followed the picks” figure — it was far too good
This page briefly showed +27.6 units from backing every pick. That number was wrong in three ways, all of them flattering. It charged no trading fees, which alone accounted for half of it. It priced every entry at the midpoint between the bid and the ask, which is not a price anyone will trade with you at — a buyer pays the ask. And it counted each strike as its own pick, so eleven different total lines on one game were counted as eleven separate wins or losses rather than one game’s outcome. Priced properly, on the picks where we recorded the real order book, backing them would have LOST about 6 units rather than made 27.
Corrected within the hour of first publishing it. The honest figure is negative, and it is the one now shown.
-
01 Sep 2026Considered adjusting the scoring model for under-predicting overs — and did not
Our probabilities came in 3.2 points below what actually happened on strikes away from the main line, which looked like the run distribution being too spread out. It is not: the market missed the same direction by 3.4 points on the same games. Both were wrong the same way, which makes it a run of high-scoring baseball rather than a flaw in our model. Changing the distribution to chase it would have been fitting noise.
A negative result, recorded because a change that was considered and rejected is as much a part of the record as one that shipped.
-
01 Sep 2026Leaning harder on the market price, because our own number has been the weaker of the two
Our published number is part betting-market price, part our own model. That split moved from 60% market / 40% ours to 80% / 20%. We re-scored every settled prediction at every possible split to find which one would have been most accurate, and the answer was 100% market — meaning our own number, on this evidence, has been making the published one worse rather than better. On its own our model scored 0.1803 for accuracy against the market price’s 0.1745, where lower is better. We moved most of the way rather than all of it, so our model keeps a real say while it earns its place back.
Fitted on 684 settled predictions from four days of games (28–31 August). The direction is statistically solid at 2.5 sigma; the exact number is not. Expect this to move again.
Standard lines vs. the tails?
Same picks, split by whether the line is the standard one or a strike out in the tail. Read the last column, not the hit rate — a high hit rate on heavy favourites can still pay very little.
| Line | Picks | Won | Units | Per pick |
|---|---|---|---|---|
| Main line only | 1665 | 61.1% | -17.80u | -0.011u |
| Every strike on the ladder | 5746 | 77.3% | -37.54u | -0.007u |
Standard lines win less often but pay roughly three times more per pick. The tail strikes look impressive on hit rate because they are mostly heavy favourites priced accordingly.
Would being pickier have paid??
If you had ignored every pick except the ones where our number was at least this far above the price, here is what you would have made. One unit per pick.
| Only picks with edge of | Picks | Won | Units | Per pick |
|---|---|---|---|---|
| 1% or more | 628 | 66.6% | +0.00u | +0.000u |
| 2% or more | 169 | 58.6% | -13.52u | -0.080u |
| 3% or more | 66 | 60.6% | +0.40u | +0.006u |
| 4% or more | 28 | 57.1% | +0.15u | +0.005u |
| 5% or more | 10 | 30.0% | -4.89u | -0.489u |
It would not have. Returns fall as the bar rises, and the top rung is negative. Read plainly: when our number disagrees most with the price, our number is usually the one that is wrong — so a bigger claimed edge is currently a warning sign, not a green light. The top rung is only 10 picks, far too few to be sure on its own, but every rung points the same way and that is harder to dismiss.
By confidence?
Does a more confident pick actually win more often? Hit rate should track the predicted column downward as you go down the table.
| Tier | Record | Hit rate | Predicted | Delta |
|---|---|---|---|---|
| Strong (65%+) | 36–14 | 72.0% | 68.6% | +3.4pp |
| Moderate (57-65%) | 103–55 | 65.2% | 60.7% | +4.5pp |
| Lean (50-57%) | 126–102 | 55.3% | 53.3% | +2.0pp |
Calibration?
A well-calibrated model’s “60–69%” bucket should win close to 65% of the time. Delta is deliberately not colour-coded — at these bucket sizes a large one is far more likely to be noise than a real miscalibration.
| Predicted | Games | Predicted avg | Actual | Delta |
|---|---|---|---|---|
| 50-59% | 296 | 54.5% | 56.1% | +1.6pp |
| 60-69% | 126 | 63.7% | 68.3% | +4.5pp |
| 70-79% | 14 | 72.7% | 92.9% | +20.2pp |
Totals & spreads accuracy?
Scored separately from the moneyline on purpose. The moneyline is a single Pythagenpat number; totals and spreads are read off a joint run distribution pinned to it, so they can be right when the moneyline is wrong and the reverse. Averaging them together would hide precisely that. Every quoted strike is scored, not just the headline line.
| Market | Settled | Main-line record | Accuracy | Our model | Market |
|---|---|---|---|---|---|
| Totals (over/under) | 4513 | 212–199 (51.6%) | 0.1700 | 0.1742 | 0.1696 |
| Spreads (run line) | 2466 | 544–278 (66.2%) | 0.1857 | 0.1879 | 0.1854 |
Main line vs. the tails?
Where a distribution assumption shows up. If the model’s average probability runs consistently above the realized rate away from the main line, its run distribution has too much spread in the tails — a real and fixable finding that a single main-line number would never surface.
| Strikes | Scored | Model avg | Market avg | Actual | Model − actual |
|---|---|---|---|---|---|
| Main line | 1233 | 41.1% | 40.2% | 41.9% | -0.8pp |
| Off the main line | 5746 | 48.3% | 47.5% | 49.9% | -1.6pp |
Recent predictions?
| Date | Game | Pick | Confidence | Final | Result |
|---|---|---|---|---|---|
| 2026-10-10 | CWS @ CLE | CLE | 55.1% | — | pending |
| 2026-10-08 | CLE @ CWS | CWS | 50.0% | 9-5 | LOSS |
| 2026-10-07 | TB @ NYY | NYY | 60.3% | 4-3 | LOSS |
| 2026-10-07 | CLE @ CWS | CWS | 53.6% | 9-3 | LOSS |
| 2026-10-07 | MIL @ SD | MIL | 51.1% | 3-1 | WIN |
| 2026-10-07 | LAD @ ATL | LAD | 56.1% | 4-1 | WIN |
| 2026-10-06 | MIL @ SD | SD | 54.9% | 3-4 | WIN |
| 2026-10-06 | LAD @ ATL | ATL | 50.2% | 3-1 | LOSS |
| 2026-10-05 | NYY @ TB | NYY | 53.1% | 2-5 | LOSS |
| 2026-10-05 | CWS @ CLE | CLE | 57.1% | 4-3 | LOSS |
| 2026-10-04 | SD @ MIL | MIL | 56.5% | 3-4 | WIN |
| 2026-10-04 | ATL @ LAD | LAD | 66.3% | 3-2 | LOSS |
| 2026-10-03 | NYY @ TB | TB | 55.9% | 0-1 | WIN |
| 2026-10-03 | SD @ MIL | MIL | 65.6% | 2-3 | WIN |
| 2026-10-03 | CWS @ CLE | CLE | 56.9% | 3-0 | LOSS |
| 2026-10-03 | ATL @ LAD | LAD | 66.9% | 3-5 | WIN |
| 2026-10-01 | PHI @ ATL | ATL | 61.6% | 2-6 | WIN |
| 2026-09-30 | BOS @ NYY | NYY | 55.1% | 2-9 | WIN |
| 2026-09-30 | CWS @ HOU | HOU | 57.6% | 7-3 | LOSS |
| 2026-09-30 | CHC @ SD | SD | 55.9% | 1-4 | WIN |
LiveEV