This article is part of the SaskPoly story. Read how it all started.
Why I Burned the Home Run Tracker and Rebuilt It
The old additive scoring model was dragging us underwater at 28%. Here's how v2 replaces astrology with interaction, spray-personalized parks, and a hard team HR gate.
Why I Burned the Home Run Tracker and Rebuilt It
A few weeks ago I looked at the backtest and felt sick.
28%. That's how often our top-3 home run picks were hitting. Not the top pick — the top three combined. I had built an additive scoring model that took a batter's base power, added the pitcher's weakness, added the park factor, added a BvP number, and spat out a rank. It felt scientific. It wasn't. It was astrology with decimals.
The #1 pick was hitting 44% of the time. The #2 and #3 picks were dragging the whole system underwater. I was swinging a dull sword and calling it sharp.
So I killed it.
What Was Broken
The old model treated every input like it mattered equally. A guy with a 12.9 base score in a +7.5 environment got the same treatment as a guy with a 6.5 base score in the same environment. The math said 12.9 + 7.5 = 20.4 and 6.5 + 7.5 = 14.0. In reality, the first guy is a completely different bet than the second. You can't add your way to truth.
BvP was the worst offender. Alex Bregman carried a 29.5 BvP and missed. Michael Conforto had 20.3 and missed. Meanwhile Gabriel Moreno hit with 14.3. The difference wasn't the number — it was the sample size. Raw BvP without a plate-appearance floor is just noise dressed up as signal.
And then there was the team problem. The model would recommend three picks in a game where the entire lineup projected 0.4 home runs. When a team hits zero, every pick dies. We were betting into dead air.
What v2 Does Differently
I spent a week studying how Ballpark Pal builds their engine. They don't rank players. They simulate. They model what the pitch should do, what the contact should do, and what the park should do — then they run the game 3,000 times and let the probabilities fall out.
I can't build that. Not yet. I don't have their data pipeline or their compute. But I can build the idea of it.
v2 is built on three principles:
1. Interaction, not addition.
The score formula is now multiplicative: (Base^0.8 × Pitch^0.6 × Env^0.9) × Recency × Platoon. A high-base hitter in a high-env game isn't just "more" — he's exponentially more. The math finally respects the physics.
2. Spray-personalized park factors.
The old model used one park number per game. v2 personalizes it by the batter's spray profile. A pull-heavy lefty at Yankee Stadium gets a different boost than an opposite-field hitter because the short porch is real and center field is deep. The "Spray Park" column in the report shows this explicitly.
3. The Team HR Gate.
If a team's expected home runs fall below 1.0, the gate kills every pick on that side. No more betting into 0.4-HR games. The report now shows Team HR Exp at the top of every matchup. When it's red, there are no picks.
4. C-Only Decomposition.
Every pick now breaks down into what the hitter did (C), what the pitcher allowed, and what the park did. You can see exactly where the score comes from. No more black box.
What the Report Looks Like Now
Instead of a flat list of #1, #2, #3, you get a probability. The Implied HR% column tells you what the model thinks the chance is. The Team HR Exp tells you whether the game is even worth looking at. The Spray column shows the batter's pull/center/oppo split so you can see why the park factor is personalized.
Some games still produce #2 and #3 picks — but only when their scores clear 8.0. Most days, you'll see #1 only. That's the point. Volume was killing us.
The Honest Truth
This is still a heuristic model. It's not a Monte Carlo simulation. I don't have 5 million pitches in a lookup table. What I have is a better way of combining the inputs I already had, plus a few new ones: a 14-day recency multiplier, a platoon penalty for same-handed matchups, and a hard BvP floor at 15 career plate appearances.
The real test starts now. I'm running v2 alongside v1 for the next week. If the #1 hit rate doesn't climb above 50%, I'll tear it down again.
That's the job. Build, measure, burn, rebuild.
If you find value in the daily briefs, consider leaving a tip. It covers hosting, data feeds, and the coffee required to stare at backtests until 3 AM.