Evidence · The claim ledger
Model performance & limits3
Every ledgered claim this site makes in this evidence area, with the sources that ground it — or, for a PAN-simulation-derived claim, the run it comes from. Source keys link back to the full reference lists on the Evidence Registry.
EmpiricalHow a model scores on data held back from its own training and how it scores at a different site are different quantitie…
How a model scores on data held back from its own training and how it scores at a different site are different quantities. The volume's disability chapter reports a named pair where the second is materially lower than the first. A figure quoted without saying which of the two it is does not tell a reader what the model will do in their setting.
Sources: wang2026a
Appears on: /pan-lab
EmpiricalThe volume's substance-use chapter reports that nearly all models in the field it reviews are developed and evaluated on…
The volume's substance-use chapter reports that nearly all models in the field it reviews are developed and evaluated on a single dataset. Where that holds, a reported ceiling describes that sample rather than a portable capability, and it should be read as the best case observed on one collection of records, not as what the tool will do elsewhere.
Sources: saba2026
Appears on: /pan-lab
EmpiricalDiscrimination statistics such as AUROC, precision, F1 and lift describe how a model separates cases on a labelled sampl…
Discrimination statistics such as AUROC, precision, F1 and lift describe how a model separates cases on a labelled sample. They are not per-interaction rates at which an error is adopted, written into a record, or corrected. The volume's research chapter treats these as different quantities, and they must never be entered into a diagram as though one substitutes for the other.
Sources: yang2026e
Appears on: /pan-lab