How to read a model's record without being dazzled
When the prediction existed, what the denominator is, whether the chances are calibrated, and why winning more often is not the same as making money.
Start with when the prediction existed
A result produced by applying a model to past races is not the same kind of evidence as a prediction recorded before the race. Both can be useful, but their design and limits must be clear.
Ask which data trained the model, which period was used to choose its settings, and which races were kept back for evaluation. If the same races keep guiding changes, they are no longer a clean test. The strongest evidence is a forward record: the prediction written down at a fixed time before racing, then settled, wins and losses alike.
Read the denominator
A claim such as "our top-rated runners keep winning" needs a denominator, a window and a selection rule written down in advance. How many runners or races were eligible? Were missing predictions excluded? Was the threshold chosen before or after looking at the results?
The complete record matters. Focusing on memorable winners tells a good story without saying anything about the method.
Compare probabilities with outcomes
Calibration asks whether the estimated chances match what happened across groups of predictions. Runners given 20% should win about one time in five across a large enough sample.
Look at each probability band, its sample size and the uncertainty around it. A good overall average can hide overconfidence in the top band or underconfidence in the bottom one.
Keep accuracy and profitability separate
Predicting the winner more often does not, on its own, make a profitable strategy. A return claim needs prices you could actually have taken, costs, availability and a defined selection rule.
Here is how our own evidence page applies this. It separates the held-out test, racing since 1 January that the model never trained on, from the forward ledger recorded at 08:45 each morning before racing. It states the race count beside every figure. And it shows the closing Betfair market beating us, because it does: on 13,018 complete races the Betfair favourite won 37.0%, our top pick 32.3% and the Timeform forecast favourite 28.7%.
Questions this piece answers
What is the difference between a backtest and a recorded prediction?
A backtest applies a model to past races after the fact. A recorded prediction is written down before the race runs. Both are useful, but only the second cannot have been shaped by knowing the result.
Is a high strike rate a sign of a good model?
Not on its own. Backing short-priced favourites gives a high strike rate and a loss. A strike rate needs its denominator, its window and the prices beside it before it means anything.
Does TrapMetrics beat the market?
No. The closing Betfair price is more accurate than every rating we have measured, ours included. What our evidence shows is that the TrapMetrics top pick wins more often than the Timeform forecast favourite, on the same races.