Primo isn't a tout picking winners by feel. It's a baseball simulator that plays every game out pitch by pitch, weighed against real sportsbook pricing under rules that are published here — including the rules that limit how much the simulator is trusted. This page explains the actual mechanics, with the real numbers from testing it. New to terms like CLV or de-vig? Check the glossary.
The core: a Monte Carlo simulation
For MLB, every plate appearance in a game is resolved individually — 20,000 times per matchup. Team-level numbers like moneyline, run line and total runs, and player-level numbers like hits, home runs and strikeouts, all fall out of the same run. That matters: pricing a team and its players separately would let the two contradict each other. Here they can't, because they come from one simulated game.
Batter and pitcher rates don't get averaged together — they combine through an odds ratio (the "log5" method). An elite strikeout pitcher facing a high-strikeout hitter produces more strikeouts than either rate alone would suggest; averaging misses that entirely. As a check: a 0.30 hitter rate against a 0.30 pitcher rate, with a 0.225 league-average baseline, correctly resolves to 0.3875 — not 0.30.
The random number generator is seeded. An unseeded simulation can't be tested, because there's no way to separate a real change in the model from ordinary noise.
Checked against real baseball
Before a simulation is trusted for anything, it has to reproduce baseball as it actually happens. Two league-average teams, run 20,000 times:
| Output | Simulated | Real MLB |
|---|---|---|
| Runs per team per game | 4.07 | ~4.5 |
| Total runs per game | 8.32 | ~8.6 |
| Home win rate | 53.9% | ~53.5% |
| Leadoff plate appearances / game | 4.86 | ~4.6 |
| Hits per game (leadoff) | 1.05 | ~1.0 |
| P(1+ hits) | 0.680 | ~0.65 |
Home-field advantage had to be added on purpose
Batting last doesn't create home advantage by itself — with two identical teams and no adjustment, the simulator returns a near-perfect coin flip (50.1%), matching research that the structural edge of batting last is worth almost nothing. Real home advantage comes from travel, familiarity and umpire tendency, none of which a plate-appearance simulation can discover on its own. The adjustment was set by directly calibrating against real MLB home win rates rather than guessed at — and across that calibration, total runs held steady around 8.33 the whole time, confirming the adjustment shifts which team scores, not how many runs get scored overall.
Regression to the mean
A bench player with 40 plate appearances and two home runs is not actually a 40-homer hitter — that's mostly noise. Every rate the model uses shrinks toward the league average, and each one shrinks at its own pace, because some stats stabilize far faster than others:
| Stat | Stabilization point (PA) |
|---|---|
| Strikeout rate | 60 |
| Walk rate | 120 |
| Hit-by-pitch | 240 |
| Home runs | 170 |
| Singles | 410 |
| Doubles | 450 |
| Triples | 1000 |
At 100 plate appearances, strikeout rate already reflects 62.5% real signal; singles reflect only 19.6%. One shared shrinkage constant for every stat would either trust singles far too early or ignore obvious strikeout ability — so there isn't one.
Singles are never read directly from the data, only derived — the stats API reports total hits, and subtracting extra-base hits is the only accurate way to isolate singles. Reading the hits number directly gives 0.1953 where the correct rate is 0.1263, a 55% overstatement that would inflate every scoring projection built on it.
And when the data simply isn't there — no confirmed lineup, no probable starter — the model returns nothing rather than a guess. A confident number built on a fictional lineup is worse than no number at all.
Blending the simulation with the market
The simulation is one independent estimate. The market is thousands of people with real money at stake. Neither gets trusted alone — and the simulation, specifically, is capped at 35% weight, deliberately below half. It has never been graded against a real result long-term, so it doesn't get to outvote the market. That cap only moves on evidence, never on confidence.
Three conditions all have to hold, multiplied together rather than averaged, so any single one failing collapses the simulation's weight toward zero:
- Completeness — the share of inputs that are real players rather than league-average fallbacks. Below 60%, the simulation isn't blended in at all. A precise simulation of a fictional lineup is still fiction.
- Precision — how much statistical noise remains at 20,000 trials versus fewer. Negligible at full trials, material if the run were cut short.
- Divergence — how far the simulation's number sits from the market's. Trust decays smoothly as the gap widens — full trust when they agree, sharply discounted once the two disagree by more than a small handful of points. A large disagreement almost always means the model is missing something the market already knows — an injury, a lineup scratch, the weather — not that the market is badly wrong.
In testing, an ideal game earned 0.327 of a possible 0.35 in simulation weight; a 25-point disagreement dropped that to 0.046. Across 361 simulation/market comparisons, the weight never once exceeded its cap.
De-vigging the market price
Every sportsbook price has the house's margin (the "vig") baked in, which has to be stripped out before two books can be compared honestly. Testing four different de-vig methods against each other turned up two findings that run against common betting folklore:
- The "proportional" method is biased toward longshots, not favorites — it strips vig in proportion to each side's size, so the favorite loses most of it in absolute terms. On a -300/+250 market it prices the favorite at 0.7241 versus 0.7363 from the method actually used here — inflating underdog edges and masking favorite ones.
- The Shin method is identical to simple additive de-vigging on two-way markets — verified to eight-plus decimal places across overrounds from 3.5% to 7.6% and favorites out to -2000. Shin only diverges from additive at three or more outcomes, and every market Primo prices is two-way, so it adds nothing here.
The default method is power de-vigging — the only one of the four that's both genuinely distinct from the others on a two-way market and biased in the direction that suppresses false positives rather than manufacturing them.
What this doesn't claim
The simulation is pre-game only. It projects a full game from first pitch and has no concept of in-play state — it runs once when lineups post, then locks. Only MLB has a simulator at all; NBA and NFL numbers here come from market pricing alone and are labeled as such, because implying a simulation ran where none did would misstate the work behind the number.
Every one of these numbers — the calibration table, the weight cap, the de-vig comparison — came from actually running the tests, not from a marketing claim. The next section is where that gets checked against reality on an ongoing basis.
The only real test of any of this is what happens after the fact — every graded pick, win or loss, published without exception.
See the track record