Completed experiment · Final results published · 11 June – 19 July 2026

Conclusion

What we learned by keeping confidence on the table

Completed experiment · 11 June – 19 July 2026 · Finalized 2026-08-01

1. The question

The World Cup floods public conversation with forecasts. Most of those forecasts are judged only by whether the pick was right. The Calibration Cup asked a narrower question: when people (and when Human Actually) stated a confidence number before kickoff, did that number deserve to be trusted after the final whistle?

2. How the experiment worked

Every prediction had two parts: an outcome and a confidence between 50% and 95%. After football-data.org marked a match final, each submission was scored with a Brier score and rolled into a Calibration Score. Trust Receipts translated the math into plain English. Visitors could keep a private browser scoreboard and, optionally, contribute to a pseudonymous public aggregate. The lab published its own group-stage slate under the same rules.

Full mechanics: methodology.

3. What happened

All 104 World Cup fixtures resolved (0 incomplete, 0 unresolved). Spain beat Argentina 1–0 (19 July 2026). Spain are champions.

Human Actually made 72 on-the-record group-stage calls. Average confidence 65%. Record 4428. Average Brier 0.2243. Calibration Score 78. The 32 knockout matches were never on the lab slate and are excluded from that total — not missing data, just unpublished calls.

The public crowd logged 131 predictions across 17 participants. Average confidence 78%. Calibration Score 67. 31 matches drew no public picks at all.

4. What the results suggest

Interpretation: stating confidence changes the cost of being wrong. The lab’s worst receipts were not spectacular upset picks — they were high confidence on favorites that drew. That pattern is familiar outside sport: the failure mode is not uncertainty, it is certainty that was never earned.

The lab’s lower average confidence and higher calibration score, next to a more confident public sample, is consistent with a simple claim: quieter numbers can be more honest. That is a pattern in this dataset, not a universal law.

5. What the results cannot prove

This was not a controlled scientific study. The public sample is small and self-selected. Human Actually did not publish knockout picks, so the lab score does not cover the whole tournament. Calibration Scores do not measure intelligence, expertise, or moral character. They measure how well stated probabilities lined up with outcomes in this experiment.

6. What I learned

Keeping the original claim visible after the result arrives is harder than building the scoring math. Narrative wants to rewrite confidence in retrospect. A receipt that still shows the number you said before kickoff is a useful friction — for fans, for labs, and for anyone shipping confident AI or decision systems.

Operationally, live tournament sync is expensive in attention and easy to leave running after the final. Closing the experiment deliberately — freezing results, stopping paid sync, leaving the evidence online — is part of treating the work as research rather than an abandoned app.

7. What I would test next

A future version would publish the full slate including knockouts, ask for draw probability more explicitly when draws are allowed, and compare calibration across stages rather than only overall. It might also separate “favorite bias” from raw Brier. The Trust Watch companion never filled with reviewed items; if revived, it should be human-curated from the start, not a research cron.

Why this still matters

The tournament is over. The site remains as a completed Human Actually experiment: methodology, historical matches, final scores, and this conclusion. The useful artifact is not a live prediction game — it is evidence that confidence can be graded, and that closing an experiment cleanly is part of the work.

View final results · Explore the matches · About this project