Deterministic by construction
Pure functions, no randomness, no clocks, canonicalized serialization. Two runs are byte-identical, which makes every result diffable, cacheable, and reproducible.
Technology
Everything interactive on this site, the evidence-cited PAN Lab and the fictional Oversight campaign, runs on a single deterministic governance engine. Same composition rules, same win oracle, same derivation traces. The Lab cites its claims; the game relaxes the evidence register for play and says so on every surface. What neither side ever relaxes is the engine.
The core
Pure functions, no randomness, no clocks, canonicalized serialization. Two runs are byte-identical, which makes every result diffable, cacheable, and reproducible.
A single verification function decides every outcome: in play, in batch judging, and in the build gate. There is no second implementation to drift out of agreement with the first.
Before a level ships, the judge enumerates its full legal move space, up to 1.5 million configurations per level, with deterministic stratified sampling above that and never silent, and proves it winnable at every difficulty tier.
A decomposed difficulty score over named, normalized factors, solution scarcity, budget tightness, trap density, time pressure, maps through an explicit model to a predicted player win rate, with per-act target bands enforced in the build gate.
The judging code carries a cryptographic lineage version recorded in a manifest. Any change to the judge fails the build until it is deliberately bumped, and a bump flags every shipped level for re-verification. Content adjusts to the judge, never the judge to a level.
Automated audits verify the prose against the mechanics: claims in explanations are replayed through the judge, cited effects must resolve to ledgered claims, and fictional content is provably isolated from the evidence channel.
How it is made
The repository enforces its own rules. Every audit ships with tests proving a clean baseline reports zero problems and that each synthetic violation fires exactly its own check: the validators are themselves validated. Evidence and fiction live in separated channels with automated leak detection in both directions. Design decisions are numbered, recorded verbatim in an append-only ledger, and superseded decisions are marked, never deleted. Content at scale is generated by a checkpointed, resumable pipeline that must pass the same gates as hand-authored content, and a batch that accepts nothing reports that honestly rather than lowering the bar.
This is what it takes to ship a trustworthy governance product with a small team at high velocity, and it is what makes the work legible to anyone who needs to evaluate it from the outside: an agency, an auditor, a partner, or an engineering organization deciding what to build on.
PAN / EMU
The Lab and the game sit on top of PAN / EMU, the 1023AI research program: a governance-network model (the Policy Actor Network) coupled to an epidemiological core (Error-Memory-User) in which AI error spreads between outputs, operators, and record stores, and either dies out or self-sustains depending on the governance around it. The research code is governed by a frozen regression contract: canonical equations, structural rules, and a pinned numerical anchor that no change may silently move. Findings are validated as direction-and-shape consistency against published external baselines, and the honest boundaries, what the model deliberately does not claim, are documented as prominently as the findings.
The research programBuilt by a small team. Governed like critical infrastructure.
Work With 1023AI