This article is part of the SaskPoly story. Read how it all started.
The Home Run Tracker v2 Now Grades Itself
The v2 report has always shown green, yellow, and red picks. Now the backtest tells you whether those colors actually mean anything. Spoiler: they do.
The Home Run Tracker v2 Now Grades Itself
When I rebuilt the Home Run Tracker last month, the report started showing every pick as green, yellow, or red. Green meant the model was confident. Red meant it was doing you a favor by showing you at all.
The backtest didn't get the memo.
It reported one flat hit rate. #1 picks, #2 picks, #3 picks — all mashed into a single number that told you nothing about whether the colors worked. A 25% overall rate could mean the model was broken, or it could mean the green picks were hitting at 40% and the red picks were dragging the average down. Those are very different problems with very different fixes.
So I fixed the backtest.
What Changed
Every pick the backtest grades now carries its level with it, using the exact same thresholds as the daily report: green at a final score of 12 or higher, yellow at 10, red below that. The counting rules didn't change — the #1 pick always counts, and #2/#3 only count when their scores clear the 8.0 show threshold.
The report now opens the aggregate metrics with a Hit Rate by Level table: picks, hits, and hit rate for each color, over the trailing seven days.
The First Week's Numbers
Here's what the last seven days of games — 98 games, 196 team-games, 277 counted picks — look like:
| Level | Picks | Hits | Hit Rate |
|---|---|---|---|
| Green | 44 | 18 | 40.9% |
| Yellow | 53 | 18 | 34.0% |
| Red | 180 | 33 | 18.3% |
Read that ordering again. Green hits at better than twice the rate of red. Yellow sits exactly where it should — between the two. The colors aren't decoration. They're the most honest thing in the report.
The uncomfortable part is the bottom row. Most counted picks are red — 180 of 277. That's structural, not a bug: the #1 pick from every live team-game gets counted no matter what, and most team-games don't produce a green. The gate does its job there — team-games with an expected HR below 1.0 get killed entirely, 0.5% of the sample last week — but within live games, the #1 pick is often a red by definition. That's why the overall rate sits near 25% while the green rate pushes 40%.
Why This Matters
A model you can't grade is a story you tell yourself. The old backtest's single hit rate made v2 look mediocre on days when the green picks were carrying everything. Now the report separates signal from volume.
The practical read for this week: if you're acting on these picks, the green tier is currently earning its confidence, yellow is a lean-not-bet zone, and red #1s are informational only. That's one week of data, not a law. I'll be watching whether the spread holds or compresses — if green and red converge toward the mean, the thresholds need retuning, and the table will tell us.
Build, measure, adjust the thresholds. That's the whole job.
You can see the graded backtest at /reports/hrt_v2_backtest/2026-09-05, and it updates every morning with the daily run.
If you find value in the daily briefs, consider leaving a tip. It covers hosting, data feeds, and the coffee required to stare at backtests until 3 AM.