Do our numbers hold up?
Every full reading Quell produces files a falsifiable prediction: the probability that the household runs out of liquid buffer within twelve months, with its uncertainty band. Once those twelve months pass, we score the prediction against what actually happened. This page is the running scorecard — including when it is too early to have one.
How scoring works
A probability can’t be “right” or “wrong” once; it can only be calibrated across many forecasts. We use the Brier score— the squared gap between the predicted probability and the observed outcome (0 = perfect, 0.25 = no better than a coin flip on a 50/50 event). Outcomes are measured against the household’s recorded liquid balances at scoring time, so the score reflects data users keep current. We publish aggregate scores only when at least 20 predictions have been scored, and the calibration table only past 50 — below that, a scorecard would be noise dressed as evidence.
The scorecard today
3 predictions currently maturing · 0 scored so far. The first cohort becomes scoreable on 2027-06-13 — twelve months after the first reading was filed. Until then there is honestly nothing to show.
What this does and doesn’t prove
Scored predictions test the model on the households that use Quell — not on the market at large, and only as well as users keep their data fresh. A household that stops updating balances scores its forecast against a snapshot. We publish the basis of every measurement alongside the scores rather than pretending the limitation away.