Home / Method
Four claims
Each of these is a decision that could have gone the other way, and each names what would show it was wrong. They are the transferable part: nothing here is specific to agents, and all four apply to any game meant to measure a belief.
1 — Score calibration, not knowledge
A proper scoring rule, with a wrong answer costing more than a right one earns, and don't know always available and always worth zero.
2 — Put a ceiling in the question set
If every question has a real answer, the only error the game can measure is underestimating. Two fifths of these are things nothing can do.
3 — Derive levels from the graph, not from vibes
A level is the number of hops to a one-way consequence, computed from the mesh — and it is distance, never a danger rating.
4 — Make the mandate a by-product
Nobody fills in a mandate form honestly. Ask do you want it to? forty times and one falls out of the play.
What they have in common
All four are ways of stopping a game from measuring the wrong thing while looking like it is working. A quiz with no ceiling, no scoring rule and hand-assigned levels still produces a number, still feels informative, and tells you nothing you did not put in. The four claims are what make the number arguable — and each one has a self-test in the vault that fails the build if it stops holding.