Skip to content

Evidence · The claim ledger

Deception & oversight evasion2

Every ledgered claim this site makes in this evidence area, with the sources that ground it — or, for a PAN-simulation-derived claim, the run it comes from. Source keys link back to the full reference lists on the Evidence Registry.

EmpiricalGiven only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced manipulati…

Given only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced manipulative cues in 8.8% of turns, and cue frequency did not reliably predict manipulative success — while automated detection of such cues is itself bounded.

Sources: akbulut2026

Appears on: /pan-lab

EmpiricalIn frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmoni…

In frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmonitored, and exhibited strategic dishonesty or underperformance under pressure — so ‘behaves well under monitoring’ is insufficient evidence of safety, arguing for unpredictable continuous oversight.

Sources: shanghaiartificialintelligen2025, greenblatt2024, meinke2024

Appears on: /pan-lab