Independent validation for sports prediction models
Every track record is a claim.
We make it a fact.
Committed before the event. Scored against the market as it stood at the decision moment.
The industry runs on records nobody can check — deleted losses, backdated wins, a screenshot for evidence. A model gets graded here against prices it actually faced, on a ledger it cannot edit afterward.
Live
§ 01The problem
Losers get deleted. Winners get backdated. We refused that.
Every model in this market is sold on a record its own author produced. There is no standard, no independent scoring, and no way to tell a real edge from a backtest fitted after the fact. The buyer's options are to take the seller's word or to run a pilot and find out slowly.
01
Ingest
Sports data pulled daily from canonical sources.
02
Train
Per-sport models, trained with cross-validation, registered with held-out metrics.
03
Predict
Calibrated probabilities, conformal sets, computed once per day.
04
Anchor
Committed to the public ledger before kickoff; every day carries a Bitcoin-attested proof.
05
Serve
Through a signed, versioned API into the mobile app; and into the public GitHub mirror anyone can run verify.py against.
No marketing. Just the receipts.
§ 02Why it stays broken
Neither side can certify the other.
Checking a record needs two things almost nobody holds: the prices a model actually faced at its decision moment, and proof the prediction existed before the event. A book cannot credibly grade its own vendor selection. A vendor cannot grade itself.
Point-in-time prices
A claim can only be checked against the market it faced — never the closing line. Without per-book history at the decision moment, contamination is undetectable, and it is the most common defect in a sports backtest.
Proof of priority
A record assembled after the fact is worthless. Establishing that a prediction existed beforehand takes a commitment made at the time, anchored somewhere the operator cannot rewrite.
Independence is not a feature here. It is the reason the thing can exist.
§ 03The substrate
A model is only checkable where the market history is.
Point-in-time prices, captured per source at open, decision time and close. That is what a prediction gets measured against — never the closing line, never a number reconstructed afterward.
- mlb
- nhl
- nba
- ncaab
- nfl
- ncaaf
Sports in scope
6
Validated — live capture
4,738
Settled and scored
996
Validated — historical replay
17,043
Settled and scored
5,167
Market types per sport
2
§ 04The protocol
We hold digests. Never picks.
A model enrols before it predicts. Each prediction is hashed and anchored before its event, revealed after settlement, then scored. By the time a pick is readable it is already history — so we cannot act on what we are validating, and you can verify that we cannot.
- 01
Register
The model enrols before it makes a single prediction. The enrolment is itself timestamped and anchored.
- 02
Commit
Predictions are hashed and anchored to Bitcoin and a public transparency log before each event. The plaintext stays with its owner.
- 03
Reveal
After settlement the predictions are revealed and checked against the commitments made before the event.
- 04
Score
Calibration, discrimination, regime control, and whether the model improved on the market's own price.
The model never leaves its owner. Only the evidence does.
§ 05Verify yourself
Don't trust us. Check us.
Three steps. Pure-stdlib Python, no toolchain. Clone the public ledger, run the verifier, check the Bitcoin attestations against the dates we claim. Every verifier release is registered and hash-checked in CI — and the sealed alpha record verifies the same way, forever.
- 01
Clone the public ledger
git clone https://github.com/SplitWinner/audit_trail - 02
Run the verifier
cd audit_trail && python3 verify.py vectors && python3 verify.py chain - 03
Hand OPERATIONS.md to your security team
open OPERATIONS.md # vendor DDQ, incident history, key ceremony
All of it works before you ever talk to us — no account, no NDA, no sales call. If we vanished tomorrow, the proof would still stand.
§ 06For desks and funds
Stop piloting claims you cannot check.
Evaluating an outside model means weeks of integration and trader attention, repeated independently by every desk, for every vendor — and at the end you still cannot tell whether the backtest was point-in-time honest. A validated record answers that before anyone writes an integration.
Design partner
One desk. One season. The standard, built with you.
We are looking for a small number of desks and funds to validate vendors alongside us while the standard settles. You name a vendor you are already considering; you get the report either way.
§ 07For model builders
Prove it without publishing it.
Enrol a model and get a record a counterparty can check — calibration, discrimination, regime controls, and provenance for every call. No feature names, no hyperparameters, no architecture. Receipts, never the recipe.
Your model stays yours
We receive hashes before the event and settled predictions after it. The model itself is never transmitted, never inspected, never stored.
A record that outlives us
The commitments are anchored in public infrastructure. If we disappeared, the proof would still stand and still verify.
Failing is allowed
A validated record is not a badge. Some models will score badly and publish anyway — that is what makes the passing ones mean something.
§ 10Questions
What we get asked.
Plain answers to the questions a skeptic asks first.
Can you guarantee wins?
No. No honest prediction service can. Sports outcomes carry real variance — anyone who guarantees is not being honest with you.
How do I know the record isn't faked?
This is the right question to ask — the industry is full of deleted losses and backdated winners. We built the system so we can't quietly delete a loss or backdate a win. Database triggers reject every UPDATE and DELETE on the predictions ledger. The day's predictions are hashed and committed to a public GitHub repo before the games start, then anchored to Bitcoin.
The full mechanism + daily reproducibility CIWhat's the difference between a prediction and a pick?
A pick is a person's opinion — easy to selectively remember, easy to forget. A prediction is a model-generated probability with a confidence score, logged and scored against the outcome. No cherry-picking, no story — just results measured over a real sample.
Aren't backtests easy to fake?
Yes. So we don't show them as live calls — we show them next to the live record, labeled, with the model's behavior in both. A backtest run on past games is honest characterisation; calling it a live call is fraud. We never blend the two.
Why is the live record so short?
Because we're new. Daily anchoring launched May 2026 — every live pick since is in the public ledger. We could pad it with our backtest history, but that's the trick we built this to prevent. The short live record is the honest signal.
See the May 2026 boundary on the verification timelineIf this works, why don't you just bet it yourselves?
We have a Kalshi entity account that a team member trades using the predictions plus their own judgment — not a mechanical bet on every pick. That's our internal modeling test, not a public-proof channel. Beyond that: real edges in sports markets are small after vig and reflexive, sportsbooks throttle sharp accounts in weeks, and posting every prediction to a public order book would telegraph the pick to anyone watching. Selling access scales without those constraints. The proof channel is the audit-trail anchor.
Interrogate it before you believe it.
Two lanes — a short note when the standard or methodology changes, or a direct line about validating a model.
Design partner
Bring us a vendor to check.
Name a model you are already evaluating and we will score it against point-in-time prices and hand you the report. No cost, no commitment, whether or not you go further.
No cost · no commitment · report either way
Read the standardsales@splitwinner.com
Operator briefing
Not ready yet? Stay wired in.
Methodology changes, incident disclosures, registry openings — the operator's notes, a few times a season. Company and role optional, but they get you routed like a desk, not a list.