MLB Betting Model Basics: Sabermetrics, Simulations, and Projection Systems

From Box Scores to Predictive Models: How Data Drives MLB Betting
My first betting «model» was a spreadsheet with pitcher ERAs and team batting averages. It was embarrassingly crude, and it lost money. But it taught me something essential: even a bad model is better than no model, because the process of building one forces you to define what you think matters and test whether you are right. Nine years later, my approach is considerably more sophisticated, but the principle is identical — let the data tell you something the market has not fully priced.
Baseball is the most model-friendly major sport. The game’s structure — discrete events, clearly defined matchups, extensive historical data — makes it amenable to statistical analysis in ways that basketball’s fluid motion or football’s small sample sizes do not. Sabermetrics, the field of baseball statistics pioneered by Bill James and popularised by the Oakland A’s front office, created the vocabulary. Betting modellers simply applied that vocabulary to the question of which team will win and by how much.
Key Inputs: Park Factors, Umpire Zones, Platoon Splits, and Weather
Every model I have built or studied starts with the same core inputs, and the sportsbooks that set opening lines use the same data, often with more granularity. The average hold rate across sportsbooks reached a record 9.7% in 2025, and that margin reflects, in part, the sophistication of the oddsmaking models you are competing against. Knowing what goes into those models is the first step toward finding where they might be wrong.
Park factors quantify how a specific ballpark affects scoring. Coors Field in Denver inflates offense by roughly 30% relative to league average, while Oracle Park in San Francisco suppresses it by around 10%. These are not guesses — they are calculated from years of game data, adjusted for the teams that play in each park. My model applies park factors to every projection: a pitcher’s expected strikeout rate at Coors is different from the same pitcher’s rate at Petco, and the difference is large enough to swing a prop line by a full strikeout.
Umpire assignment matters more than most casual bettors realize. Home-plate umpires have measurably different strike zones. Some call a wide zone that benefits pitchers, generating more called strikes and fewer walks. Others run a tight zone that puts pitchers behind in counts and increases offense. I maintain a database of umpire tendencies — zone width, called-strike rate, walk rate — and adjust my projections by 2% to 4% depending on the assignment. It is a small edge, but small edges compound over 162 games.
Platoon splits capture the difference in performance between left-handed and right-handed matchups. A left-handed batter facing a left-handed pitcher typically produces 15% to 20% less offense than the same batter facing a right-hander. Lineup construction exploits this, and so should your model. When a team stacks right-handed hitters against a left-handed starter, the projected run total should reflect the platoon advantage — and the sportsbook’s line usually does, but not always perfectly.
Weather is the wild card. Temperature, humidity, wind speed, and wind direction all affect ball flight. A 6-degree Celsius temperature increase adds roughly 1.2 metres to a fly ball’s carry, which can turn a warning-track out into a home run. I pull weather data two hours before first pitch and adjust my total projections accordingly. The adjustment is modest — usually half a run in either direction — but on a total line of 8.5, half a run is the difference between over and under.
Projection Systems and How They Generate Prop Lines
Projection systems are models that forecast player performance based on historical data, aging curves, and contextual factors. The most widely referenced public systems in baseball are ZiPS, Steamer, and THE BAT X. Each uses slightly different methodologies — weighting recent performance versus career data, adjusting for park and league effects, incorporating minor-league data for young players — but they converge on similar outputs: projected batting average, home run rate, strikeout rate, and other core metrics.
For betting purposes, projection systems serve as the foundation for generating fair lines on player props. If ZiPS projects a pitcher to average 7.1 strikeouts per start against an average lineup, and the sportsbook posts a line of 6.5, the model suggests the over has value. The gap between the projection and the line is where the edge potentially lives, though the word «potentially» does important work — projections are estimates, not prophecies.
I use a composite approach, blending outputs from multiple projection systems rather than relying on any single one. The blend smooths out idiosyncratic assumptions and produces a more stable estimate. When three systems agree that a hitter’s projected home run rate is significantly above the implied rate in the sportsbook’s prop line, I have more confidence in the bet than when one system says yes and two say no.
Monte Carlo Simulations and Win Probability Outputs
If projection systems tell you what should happen on average, Monte Carlo simulations tell you the range of what could happen. The concept is straightforward: simulate the game thousands of times, each time introducing random variation around the projected averages, and observe the distribution of outcomes. A Monte Carlo simulation of a single MLB game might run 10,000 iterations, producing a win probability, a distribution of total runs, and a distribution of individual player statistics.
I started running my own simulations about four years ago, and the insight that changed my approach was this: the average outcome is far less useful than the distribution. Two games might both have a projected total of 8.5 runs, but one game’s distribution is tightly clustered around 8-9 runs while the other has a wide spread from 4 to 14. The flat total line is the same, but the betting implications are different. The wide-distribution game offers more value on overs and unders at extreme prices, because the tail outcomes are more likely.
Building a simulation from scratch requires programming ability and access to player-level data, but you do not need to build one from scratch to benefit from the approach. Several public platforms publish simulation-based win probabilities and prop projections. The value for a bettor lies in comparing those simulation outputs to the sportsbook’s lines and identifying divergences. When a simulation says a team has a 58% chance of winning and the sportsbook’s line implies 52%, the 6-point gap is a signal worth investigating.
No model is perfect. The best models in baseball still lose roughly 45% of their bets. The advantage is measured in percentage points, not in certainties, and it requires volume and discipline to realize. What modeling gives you is a framework for making decisions consistently, evaluating those decisions honestly, and improving over time. That is what separates a bettor with a process from a bettor with a feeling — and over a full MLB season, the process wins. For more on how to translate model outputs into actionable prop bets, the player props strategy guide applies these concepts to specific markets.
What data inputs do MLB betting models typically use?
Core inputs include park factors, starting pitcher projections, lineup construction, platoon splits, umpire strike-zone tendencies, weather conditions, bullpen availability, and recent form. Advanced models also incorporate pitch-level data, spray charts, and catcher framing metrics. The goal is to generate a probability estimate for each game outcome and player performance line that can be compared to sportsbook odds.
How accurate are simulation-based MLB betting projections?
The best publicly available MLB models win between 54% and 56% of bets on moneylines and totals, which is enough to generate long-term profit after accounting for the vig. No model consistently exceeds 60% accuracy over a full season. The value of simulations lies not in predicting individual outcomes perfectly but in identifying systematic mispricings across a large volume of bets where the model’s edge compounds over time.
Escrito por los editores de «mlb Players Betting».