10 to the 23 AI logo

Domain Atlas / Housing & homelessness services

Case fileGreater London, United Kingdom (all 33 London local authorities — 32 boroughs plus the City of London — with the pan-London bodies GLA and London Councils)large deployment

London's Strategic Insights Tool: one linked memory of rough sleeping, read by every borough

London's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately governed systems - CHAIN street-outreach contacts, In-Form charity casework, and H-CLIC borough statutory applications - into a single rough-sleeping journey per person that is read, in aggregate form only, across all 33 London local authorities; the tool makes no individual-level determinations, and after the build vendor's data-processor contract ended on 2 February 2024 the Greater London Authority contracted Homeless Link, which also operates the CHAIN source system, for its ongoing hosting, management, and maintenance.[3]

What happened

The Strategic Insights Tool for Rough Sleeping (SIT) is a pan-London data service that links records from three separately governed systems into a single rough-sleeping "journey" per person, and lets commissioning teams across all of London read the result. The three sources are CHAIN (street-outreach contacts, commissioned by the Greater London Authority and provided by the homelessness charity Homeless Link), In-Form (charity accommodation and hostel casework), and H-CLIC (borough statutory homelessness applications). The linkage is done with Splink, an open-source probabilistic-matching library, on identifiers including fuzzy-matched names, National Insurance numbers, dates of birth, and phone numbers. It is important to be precise about what the tool is and is not: it makes no individual-level decisions and runs no risk score. Users see aggregate, population-level views only (plus records their own organisation uploaded), and the project team states plainly that it "is not a substitute for published data and reports." Its purpose is to inform cohort-level commissioning, strategy, and funding — not any decision about any person.

The project was initiated in 2022 by the Life off the Streets Partnership (the GLA and London Councils, with advisory support from Bloomberg Associates). The applied-AI consultancy Faculty was appointed technical delivery partner and sole data processor on 5 June 2023; a minimum viable product went live on 8 September 2023 in a pilot with four boroughs (Camden, Hillingdon, Lambeth, and Westminster) and their service providers. An information-governance review of that pilot phase, completed 31 October 2023, cleared the wider Phase 2 rollout, which reached all London boroughs plus service providers on 29 February 2024. Faculty's build contract ended on 2 February 2024, after which the GLA contracted Homeless Link — which already operates the CHAIN source system — for ongoing hosting, management, and maintenance. By March 2025, LOTI reported 45 organisations contributing data, with the service hosted on AWS microservices.

The tool's documented weakness is a matter of arithmetic the project has been unusually candid about. The matcher accepts an association only above an 85% probability threshold, chosen to minimise false positives. The project's own Phase 2 Data Protection Impact Assessment (version 2.0, dated 23 October 2023 and published on the LOTI site in April 2025) reports 91% recall and states the risk in plain terms: "we miss 9/100 matches and numbers subsequently appear lower in places where they should be higher," adding that recall will vary as new data of varying quality is ingested. No false-positive rate is published, and these figures are self-reported by the delivery team rather than established by an independent audit. Because every borough reads the same consolidated layer, that undercount is not one team's local error — it is inherited across the whole city at once. The write cadences that feed the layer are heterogeneous by design: an automated pipeline updates CHAIN weekly, service providers submit In-Form monthly by manual upload or API, and boroughs submit H-CLIC quarterly aligned to their government (DLUHC) returns, with history from 1 January 2022. When linked records conflict, a prioritisation step selects the most reliable source; pre-matching cleaning standardises phone and National Insurance numbers, lowercases names, strips suspicious names, and removes duplicates.

The governance surrounding the tool is a strongly documented, multi-actor structure — which is exactly what makes the one unmeasured surface stand out. All participating boroughs and charities, together with the GLA and London Councils, are joint data controllers under a pan-London Data Sharing Agreement, with the per-organisation compliance position set out in the DPIA. During user testing, a reader discovered that stacking filters could reduce a displayed insight to one or two individuals; the team assessed the re-identification risk and imposed ONS-style small-cell suppression, so any visualisation output of five or fewer now displays as "equal to or less than 5." Retention is five years on a per-individual basis, aligned to the way the DLUHC rough-sleeping indicators are defined (someone not seen sleeping rough for five years stops being treated as an existing rough sleeper), and unmatched service-provider data is deleted automatically. A phase-1 data-minimisation review found that several sensitive shared fields — including substance misuse, current mental health concerns, pregnancy status, prison history, care-leaver history, and entitlement to welfare benefits — went unused in the MVP; notably, the decision was to retain and keep collecting them for planned future features rather than to delete them, so that unused sensitive data "is currently retained in the SIT environment but is not released or visible to users in any way." What was removed was data submitted outside the scope of the request, and the unmatched records; meanwhile CHAIN time-and-location granularity and H-CLIC eligibility, priority-need, and duty-type fields were expanded to build the journeys.

The evidence for the tool's effect is thinner than the evidence for its design, and its provenance should be kept explicit. Faculty's own case study reports that 40 organisations have supplied data with 151 onboarded users (its headline stats cite 32 London boroughs onboarded and 12 service providers), and describes the work as giving "an understanding of the rough sleeper population for the first time" and replacing manual data tasks — all vendor claims. The GLA Chief Digital Officer's first-year retrospective (December 2024) reports, qualitatively, that the tool helped commissioners "challenge assumptions about local rough sleeping patterns" and revealed "previously hidden connections between street homelessness and Housing Options services," and names the funders (London Housing Directors, the GLA, and central government via MHCLG), with delivery support from the homelessness social enterprise Beam. There are no published metrics, counterfactual, or independent evaluation behind those first-year claims. The scale of the underlying problem is set out in two mayoral decisions: MD3161 (23 August 2023) approved £144,535 of London Councils funding for CHAIN staffing to support the SIT, within a £1,401,011 package, and recorded 10,053 people seen sleeping rough in London in 2022-23 (a 21% year-on-year increase); MD3331 (4 February 2025) approved £736,000 to expand CHAIN capacity through Homeless Link for 2025-26 to 2027-28 and recorded 11,993 people seen sleeping rough in 2023-24, a 19% increase — the rising population whose records the SIT consolidates, though that decision does not itself name the tool. Stated future ambitions — predictive demand-forecasting and the integration of further datasets such as probation, health and care, and evictions — remain aspirations in every source, not deployed features.

The sociotechnical reading

Most of the Atlas's housing systems put a number on a person: a prioritization score that decides who reaches scarce help first, a prevention model that ranks who to reach, a risk tier that routes a case. The Strategic Insights Tool is a different object, and it teaches a different lesson — one about consolidated memory. It scores no one and decides nothing about any individual. What it does is take three separately governed record systems, each with its own custodian and its own update clock, and probabilistically link them into a single shared picture of who is sleeping rough — and then let all 33 London local authorities read that one picture. On the system map, that is a very specific shape: not a model whose output flows to a decision, but a memory layer that many independent readers consolidate onto at once. And the governance question that shape raises is not "is the score fair?" It is "when everyone reads the same linked memory, whose error do they all inherit?"

The answer, by the tool's own admission, is a specific and quantified one. The matcher is tuned conservatively — it accepts a link only above an 85% probability threshold to avoid false positives — and it reaches 91% recall, which its DPIA is candid enough to spell out as roughly 9 in 100 true cross-system matches missed, so that "numbers subsequently appear lower in places where they should be higher." In a per-borough system, that undercount would be a local error, contained within one organisation's records, discoverable and correctable there. Consolidated into one shared layer, it becomes something else: a single error structure inherited simultaneously by every reader, a correlated blind spot across the whole city. There is no second, differently sourced view to disagree with it, because the whole point was to build one view. That is the distinctive hazard of consolidation — it does not create the error, but it turns one system's error into everyone's, all at once, and removes the diversity of perspective that would otherwise surface it as disagreement.

What makes the case instructive rather than alarming is that the error is known and conceded — and that almost everything around it is unusually well governed. This is a project with a published DPIA, a re-identification finding caught in user testing and answered with small-cell suppression, a minimum-data-set review, retention rules tied to a statutory definition, and a joint-controller structure spanning every borough. The one surface that is not held to the same standard is the accuracy of the linkage everyone consolidates onto: the recall is self-reported by the delivery team, no false-positive rate is published, it varies with each ingestion's data quality, and no routine reconciliation checks the linked layer back against its sources — nor does any independent evaluation test whether the shared figures actually improved commissioning. So the governance that matters for a node like this is not the governance the rest of the domain reaches for. There is no risk score to make fair, no eligibility gate to appeal, no better model that changes the shape of the problem — a sharper matcher still produces one shared error inherited city-wide. The instruments that fit are the ones for a shared memory: reconcile the linked copy against its sources on a rhythm, because recall drifts; carry the known undercount onto every read, so a commissioner sees an estimate with a direction of error rather than a headcount; and stand up an independent look at whether the consolidated picture helped, so the shared source keeps a check outside itself. The distinct lesson the Atlas draws here: when many independent readers consolidate onto one shared memory, a single linkage error stops being local and becomes everyone's blind spot at once — and the governance that then matters is not the matcher's accuracy but whether that shared layer's known error is reconciled and made legible, or silently inherited as ground truth. A picture everyone trusts needs a check no single reader can supply. The honest boundary throughout: served people — people sleeping rough, and the services they do or do not receive — are not modeled in the paired Lab; the harm surface is aggregate-statistical (a systematic undercount, a re-identification risk, a commissioning picture treated as more authoritative than it is), the accuracy and adoption figures are self-reported or vendor-attributed rather than independently audited, and this was never an individual-scoring system — any reading that implies case-level adjudication misreads it.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

faculty2024GroundingVendorSave

Faculty, Improving insights into homelessness in London with AI (vendor case study, c. 2024) https://faculty.ai/ourwork/loti

https://faculty.ai/ourwork/loti

Grounds: model org: london_rough_sleeping_sit

greaterlondonauthority2025GroundingGovernmentSave

Greater London Authority, MD3331 Rough sleeping funding and services 2024-25 to 2027-28 (2025) https://www.london.gov.uk/who-we-are/governance-and-spending/promoting-good-governance/decision-making/mayoral-decisions/md3331-rough-sleeping-funding-and-services-2024-25-2027-28

https://www.london.gov.uk/who-we-are/governance-and-spending/promoting-good-governance/decision-making/mayoral-decisions/md3331-rough-sleeping-funding-and-services-2024-25-2027-28

Grounds: model org: london_rough_sleeping_sit

Topics: ai-governance

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalLondon's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately gove…

London's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately governed systems - CHAIN street-outreach contacts, In-Form charity casework, and H-CLIC borough statutory applications - into a single rough-sleeping journey per person that is read, in aggregate form only, across all 33 London local authorities; the tool makes no individual-level determinations, and after the build vendor's data-processor contract ended on 2 February 2024 the Greater London Authority contracted Homeless Link, which also operates the CHAIN source system, for its ongoing hosting, management, and maintenance.

EmpiricalThe Strategic Insights Tool's matcher accepts an association only above an 85% probability threshold chosen to…

The Strategic Insights Tool's matcher accepts an association only above an 85% probability threshold chosen to minimise false positives, and the project's own Phase 2 Data Protection Impact Assessment reports 91% recall - conceding that roughly 9 in 100 true cross-system matches are missed so that 'numbers subsequently appear lower in places where they should be higher' and that recall varies as new data of varying quality is ingested; no false-positive rate is published, the accuracy figures are self-reported by the delivery team, and no independent evaluation of the tool's decision impact exists.