Domain Atlas / Housing & homelessness services
Calgary Drop-In Centre: interpretable screening a shelter's own staff choose to check
Work with this case in the PAN Lab ↗
At the Calgary Drop-In Centre, a University of Calgary engineering group and the NGO shelter operator built deliberately interpretable screening for chronic and episodic shelter use - explicit stay-count thresholds (for example 81 or more stays in a 90-day window) and database-queryable rules derived from the shelter's own administrative records, reported to flag candidate clients at a median of about 98 days versus 285 days under the Government of Canada definition and 365 under the Alberta definition - and, rather than surface a risk score, deployed a co-designed data-navigation interface that shows frontline staff raw client histories; no fetched source confirms the thresholds running as an automated production screener, and the deployed, studied artifact is the raw-history interface.[3]
What happened
The Calgary Drop-In Centre is a registered charity and nationally accredited nonprofit that connects adults experiencing homelessness with shelter, housing, and health programs from its downtown Calgary location — an NGO shelter operator, not a municipality, and described in the research literature as one of the largest emergency shelters in Calgary (500+ beds, and by the most recent study nearly 7,000 unique clients a year, 6,839 in 2022-23). Over roughly five years it partnered with a University of Calgary engineering research group (the Messier group at the Schulich School of Engineering) on a two-part project: deliberately interpretable screening for chronic and episodic shelter use, and a co-designed interface that surfaces that data to frontline staff. It is important to be precise about what was built. The screening is a set of explicit, human-readable rules, not an opaque risk score, and the deployed, studied artifact is a data-navigation interface that shows raw client histories rather than a recommendation.
On the model side, researchers analyzed 5,431,521 shelter entries for 34,577 unique clients (July 2007 to January 2020; 18,398 after inclusion criteria) to derive simple threshold tests: RAPID-Chronic flags a client with 81 or more stays in a 90-day window, and RAPID-Episodic flags 2 or more episodes of shelter access in 90 days. The tests identify candidate clients at a median of about 98 days — versus 285 days under the Government of Canada chronic-homelessness definition and 365 days under the Government of Alberta definition, roughly 187 and 267 days earlier respectively — while, in the paper's modelling, averaging 194.8 stays saved and 874.4 days of shelter-tenure reduction per referral. A follow-up rule-search framework using the OPUS algorithm over 12 years of the shelter's data (5,060,302 interactions for 41,935 profiles; 3,191, or 9.9%, of the 32,346 clients retained after inclusion criteria met the Canadian chronic criteria) cut median at-risk identification from 297 to 162 days and produced database-queryable rules — for example, one flagging clients with many sleep entries but few barring events, at about 85% recall and 60% precision — characterizing chronic users who "fly under the radar" with high shelter use but few bar or counselling events. A deliberate comparison of the thresholds against logistic regression and neural networks found the machine-learning methods scored better on conventional metrics but selected cohorts with very similar characteristics, and that false positives were often still good housing candidates — grounding the choice of simple, interpretable techniques for a resource-limited nonprofit. The rule-search work explicitly frames using administrative data as a way to avoid a re-traumatizing intake vulnerability survey.
The deployed artifact, however, is not an automated threshold screener — a distinction the sources are careful about. Over five months in 2022, the team co-designed and deployed a data-navigation interface with 11 frontline staff (senior managers, shelter managers, shelter coordinators, security staff, impact analysts, and housing-placement staff); by the CHI 2023 publication its second iteration was in live use with client data during the shelter's weekly Bar Review Committee meetings. The October 2025 study names the tool the Bar Review Data-navigation Interface (BRDI, pronounced "birdy"): a Lookup page for filtering barred clients and a Deep Dive page showing demographics, housing status, logs, active and inactive bars, and a check-in column chart with barring history superimposed — raw histories, not a risk score. No fetched source confirms the RAPID thresholds or OPUS rules running as an operational scoring pipeline; the deployed, studied object is the interface.
The distinctive contribution is what happened when the team studied its own staff using it. Based on 2022 to 2024 fieldwork — 16 staff across 7 role categories, 29.5 hours of interview and observation data, five Bar Review Committee observations, three co-design sessions, and three deployed versions — staff articulated what the researchers call a "data-outsourcing continuum." They were reluctant to outsource high-stakes barring decisions, treating the data as "a starting point for collaborative discussions" and, in one staff member's words, as the canvas: "the technology is the canvas, and then we kind of do the painting." For lower-stakes housing-triage decisions, they reported more willingness to accept automated, data-driven recommendations. Their resistance to abstraction was explicit — "situations…are very complex. We'll need to read everything in order to get the big picture." The interface remained in active use at the study's end, and staff had initiated a follow-on dashboard project with the shelter's IT team. The evidence for the design is peer-reviewed and quantitative; the evidence for the deployment's effect is qualitative and, importantly, authored entirely by the embedded research team. There is no independent evaluation or audit, and no usage logs, override or agreement rates, or decision volumes are published; the 2025 deployment study is a preprint with no independent venue or DOI found as of mid-2026. The work ran under University of Calgary Conjoint Faculties Research Ethics Board approval (including REB21-1121) and was supported by NSERC, the Government of Alberta, the shelter itself, and Making the Shift; the lab's stated stance is to "provide information to support the human staff doing front line work" rather than to automate care decisions.
The sociotechnical reading
Almost every predictive tool in this Atlas is judged on whether its users trust it too much — automation bias, a habit of deferring to the score, a floor that becomes load-bearing before anyone notices it shift. Calgary is the case that measured the deference directly, and found it was neither uniform nor accidental. Over two and a half years embedded in the shelter, the researchers watched their own frontline staff use the tool and documented a stakes-dependent gradient: for a high-stakes barring decision the staff read everything and treated the data as "the canvas" they then paint on; for lower-stakes housing triage they were more willing to let the data recommend. That gradient is not a failure to govern — it is governance the design earned. By choosing explicit, interpretable rules over an opaque score, and by building an interface that shows raw client histories rather than a recommendation, the builders gave their users a tool they could audit instead of defer to, and the users calibrated their trust to what was at stake. This is the Atlas's clearest evidence that vigilance can be designed into a tool's form, not merely trained into its users — and it is the reason the case sits apart from the sibling municipal risk-score tool, which is an opaque classifier a caseworker can only trust or override, where Calgary is a set of legible rules and a raw-history interface a caseworker can read against.
But designing vigilance in does not remove the risk; it relocates it. When a score can no longer be over-trusted — because there is no score — the danger migrates to two quieter places, and the map shows exactly where. The first is the low-stakes tail of the deference gradient. The very calibration that makes the tool safe for barring decisions is, by the staff's own account, deliberately relaxed for triage, and that relaxed end is where an unwritten habit is thinnest and where new staff, who did not spend years learning to treat the canvas as incomplete, are most likely to accept a data-driven view as the whole story. The levers that fit are not another model upgrade but the ones that keep a practice alive: a vigilant channel that holds the low-stakes reflex, and protection for the skill of reading the record rather than accepting its summary. The second, and more structural, is the closed recording loop. The tool writes nothing — only humans do — so the loop runs from the staff who author the barring and log records, into the store, into the screening that reads them back, and out to the staff again. That is why the "fly under the radar" finding is so telling: those clients show up plainly in the routine attendance counts but almost nothing in the discretionary log, because their situation never generated an entry, so the record's sparsity for them is a product of how staff record, not of what is true. A raw history read as complete quietly drops them — and the interface, by design, presents it as complete. The map's answer is a check the users cannot supply from inside their own practice: reconcile the discretionary log against the objective attendance record to surface whom the log omits, mark on every read what the record does and does not capture, and — because everything known about this deployment was written by the team inside it — put an independent evaluation on a rhythm from outside. The distinct lesson the Atlas draws here: a tool can be built so its own users audit it rather than defer to it, and when it is, the work of governance is not to make the model more accurate but to defend the two things the design leaves exposed — the low-stakes end of a calibrated deference, and the records the users themselves author and then read back as ground truth. A tool built to be audited by its users still needs a check its users cannot give it. The honest boundary throughout: served people — adults experiencing homelessness, and the shelter, barring, and housing outcomes they do or do not receive — are not modeled in the paired Lab; the harm surface is institutional (an omission in the record, a deference gradient, a loop read back as truth), the deployment evidence is participant-researcher-authored rather than independently audited, and this was never an automated scoring system — any reading that implies a risk score or a determination misreads it.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.