A game makes you say what you think — before you are told
Everything else we publish can be nodded at. A graph, a standard, a report, a map: you read it, you agree, and neither of us learns whether you already knew. A game does not allow that. It makes you commit to an answer first, and the gap between what you said and what is true is data that could not have been collected any other way.
That is the argument this site exists to make, and it is falsifiable: if play data shows people are as well calibrated as they think they are, the games measured nothing worth measuring. We publish the scores, so you will be able to tell.
The subject is what your agent can actually do
The first games ask one question in two directions. You have given an AI agent access to something — a machine, a repository, a mailbox, a cloud account. Can it do X? And, separately, do you want it to? Answer that forty times and you have written a draft mandate without meaning to, and the delta between what the agent can do and what you wanted it to do is the thing nobody has written down.
That delta is the whole subject of RiskMandate — the business risk layer for autonomous systems — which is the project these games are part of. RiskMandate governs the right to act: what an agent may do, granted by whom, for how long. A form would ask you for your half of that and you would answer aspirationally. A game gets it out of you as a by-product of playing.
What happens to a delta once you have one is RiskMandate's answer rather than this site's, and the game that produces one carries it: what to do next. Here, the relevant question is the narrower one — why a game gets it out of you at all.
Play the first one
That is the real game, running out of the encrypted vault it is published in — no copy of it exists on this site. It has its own player-facing home at what-can-it-do.games.sgit.ai, which is the link to send someone who just wants to play. This page is for the reader who wants to know why it exists.
The games
What Can It Do? — the scoreboard
Name your agent and where you run it. The board asks, capability by capability, can it? and do you want it to? You score for calibration, not for knowledge, and the end screen hands you the mandate you assembled without meaning to.
Which Agent Is It? — the floor plan
Think of an agent. Cheap questions narrow the field while you chalk which wings of the building you think it can enter. Then the doors open and the prediction gap is shown per capability.
Ideas & feedback — the reply channel
Not a game: the thing that turns one into a conversation. What players argue with becomes ideas, grouped into themes, answered with a published position — as a graph, in the same shape as the mesh the questions come from.
Why a maturity label on every one
Games arrive half-built and stay that way for a while, and a catalogue that hides that is a catalogue nobody can use. Every game here carries one of five rungs, and each rung has a test that moves it up rather than a feeling: playable means a stranger finished a run without being told how; measured means a release note cites play data. The ladder, stated once so a label means the same thing on every card.
What this site does not claim
- That a good score is safety. Calibration is a property of the player, not of the environment they are describing. Somebody perfectly calibrated about a badly configured agent is still running a badly configured agent.
- That the board is right about your setup. Both games measure you against a published profile — what a vendor says a product does — which is the weakest tier of evidence there is. It is not an audit of your deployment and never claims to be.
- That the points mean anything yet. The values and cut-points are arbitrary until there is enough play data to fit them. That is stated inside the game too.
- That the mandate is a mandate. It is what one person said while playing, badged in the game itself as a draft you wrote while playing.
cf04d8a9bac6…7b505f:pg87npy3 — open it read-only in a new tab. The games vault, with Which Agent Is It? and the 9 September version of both on branch release-2026-09-09: f94c8b1d4235…111118:4evnlwrj — open it; the full string is on its page at sgit.ai. Both are read keys: they cannot write, which is what makes publishing them safe.