The standard
One record format.Every model, the same questions.
A validated record answers six things in the same order, whoever built the model and whatever it predicts. No bespoke methodology per vendor, no metric chosen after the result is known. This page is the whole rubric — read it before you read any record scored against it.
§ 01What every record reports
Six questions, asked identically.
Each one exists because a model can pass the others and still be worthless. Together they are hard to game; individually every one of them has been gamed.
- 01
Point-in-time correctness
The evaluation reads what the model could have known at its decision moment and nothing later. Prices come from per-book history captured at or before that instant, never walked forward. This is the check most sports backtests fail, and failing it inflates every downstream number.
- 02
Calibration
When a model says seventy percent, does the thing happen seventy percent of the time. Reported as reliability, resolution and expected calibration error — because a model can be perfectly calibrated and carry no information at all.
- 03
Discrimination
Does a higher stated probability actually correspond to a more likely outcome. Calibration without discrimination describes a model that is honestly useless; both are reported, neither alone.
- 04
Regime control
An apparent edge is usually a market condition. Every record is permutation-tested against a price-matched baseline drawn from the same pool, so a result that merely rode a favourable stretch is separated from a result that selected within it.
- 05
Additivity
Does the model improve on the price the market already published. This is the question a desk actually needs answered, and it is the one a self-reported track record never addresses.
- 06
Provenance
Every prediction carries a commitment made before its event, anchored where the operator cannot rewrite it. Without this the other five are computed over a record that could have been assembled afterward.
§ 02What a record never contains
Receipts, never the recipe.
A validated record carries evidence about a model, never the model. No feature names, no hyperparameters, no architecture, no training data lineage. A vendor can hand a buyer the record without handing them the thing that produced it — which is the only arrangement under which anyone would enrol.
The evidence travels. The model stays home.
§ 03The shape of a record
What comes out the other end.
Field names, not fabricated figures — a real record renders from measured values and carries the sample size beside every one of them.
| source | per-book price at decision time |
|---|---|
| scope | predictions · settled |
| calibration | reliability · resolution · expected error |
| discrimination | AUC against realised outcomes |
| regime control | permutation vs price-matched baseline |
| additivity | improvement over the market price |
| provenance | committed pre-event · dual-root anchored |
Illustrative field layout. Values are populated from the scored record, never authored.
§ 04Failing is a result
A verdict is allowed to be no.
Most models that get scored will not clear the additivity bar. Records that fail are published in the same format as records that pass, because a registry where everyone passes is a directory, not a standard.
A validator that cannot fail anyone is an advertisement.
Bring us a model to check.
Enrol one of your own, or name a vendor you are already evaluating. Either way you get a record scored against point-in-time prices, in the same format as every other record in the registry.
No cost · no commitment · report either way