10 to the 23 AI logo

Domain Atlas / Industrial QA & operations AI

Case fileRegulated pharmaceutical manufacturing (GxP quality process; a regulator's developing AI framework)large deployment

Tuned to over-reject and the miss it still has to guard

Explore this deployment in the PAN Lab ↗

Machine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects that manual inspection or fixed-rule cameras would otherwise judge. In this safety-critical, regulated manufacturing setting the error trade-off is asymmetric and deliberate: a false accept — a missed defect in an injectable that reaches a patient — is a patient-safety failure, while a false reject — scrapping a good vial — is a cost, so the system is tuned to over-reject rather than risk a miss. Because the inspection sits inside a validated pharmaceutical quality process, the AI cannot simply be switched on; it must be qualified within that process, and a regulator is actively developing the framework for how AI in drug manufacturing should be validated and monitored.[2]

What happened

A pharmaceutical manufacturer uses machine-learning automated visual inspection to examine filled injectable drug products — vials and syringes — for particulate contamination and cosmetic defects, a task previously done by human inspectors and fixed-rule camera systems. The deployment is documented in the pharmaceutical quality literature, and it sits inside a regulated, safety-critical manufacturing process, which changes what the governance is about.

The defining feature is the asymmetry of the error trade-off, and that it is chosen on purpose. The two errors an inspection can make are not equivalent here. A false accept — a genuine defect that the system passes, so a compromised injectable reaches a patient — is a patient-safety failure, the thing the entire inspection exists to prevent. A false reject — a good vial the system scraps as defective — is a cost: wasted product, lower yield, but no patient harmed. Because these are so unequal, the system is deliberately tuned to over-reject: to err toward scrapping good product rather than toward passing a bad unit. That is the right call for a safety-critical line, and it is a governance decision, not a neutral default — someone chose which error to make, and the choice is defensible precisely because it is owned.

Putting AI inside a validated quality process is the second governing fact. In regulated pharmaceutical manufacturing, a change to an inspection step is not simply deployed; it must be qualified — demonstrated, within the validated process, to perform as intended and to keep performing. For a fixed-rule camera this is well-trodden. For a machine-learning model whose behavior can shift with the product mix, the sensors, or a retraining, the qualification question is harder, and the regulator's framework for how AI specifically should be validated and monitored in drug manufacturing is still being developed. So the qualification of the model's behavior over time is an emerging check, not a settled one — the deployment is inside a mature quality system, using an AI the quality system does not yet have a finished way to qualify.

The over-reject tuning lowers the visible risk but does not remove it, and this is where the human backstop matters. Tuning toward false rejects makes a missed defect rarer, not impossible; some defects will still slip, and the human inspector is the layer that is supposed to catch them. That backstop is exactly what over-trust erodes. If inspectors come to treat the AI's pass as authoritative — deferring to it, scrutinizing less because the machine has already looked — then the residual misses the tuning was designed to guard against pass through the one layer that was there to catch them. The failure mode is quiet and precisely located: it lands on the missed defect the whole asymmetric design was built to prevent, at the moment the human stopped being an independent check.

The honest reading is that this is a well-governed use of AI in a safety-critical setting — the error trade-off is chosen deliberately in the safe direction, and the deployment respects a regulated qualification process — with two live governance surfaces that the maturity of the surrounding quality system does not automatically close: the AI-specific qualification the regulator is still framing, and the human backstop that over-trust can hollow out. The residual risk after over-reject tuning is small and real, and it is covered only if both surfaces hold.

The sociotechnical reading

This case closes the industrial-QA domain in a regulated, safety-critical setting, and it makes the domain's error-trade-off point in its sharpest form. In-line inspection anywhere involves choosing which error to make; here the choice is life-or-death asymmetric — a false accept can harm a patient, a false reject only costs product — so the system is deliberately tuned to over-reject. The map reads this as the error trade-off owned correctly: not a classifier (a model that sorts each case into a category) sitting at whatever threshold, but a governance decision made in the safe direction on purpose. The instruction is that the trade-off is always a choice, and in a safety-critical line the choice should be explicit, owned, and biased toward the survivable error.

Two surfaces remain live even when the trade-off is chosen well, and the case draws both as latent checks. The first is qualification: an AI inside a validated quality process must be qualified and monitored for drift, and because the regulator's AI-specific framework is still developing, that qualification is an emerging check rather than a finished one — the deployment uses an AI the mature quality system does not yet have a settled way to validate over time. The second is the human backstop: over-reject tuning makes a miss rarer, not impossible, and the human inspector is the layer meant to catch the residual misses. Over-trust erodes that layer precisely where it is load-bearing, so the quiet failure lands on the exact defect the asymmetric design was built to prevent.

The generalizable lesson across this domain's three orgs is now complete. The benefit runs through the human loop (BMW's response loop; heavy industry's feedback loop), the failures are the loop's calibration (whether a score of, say, 30% really does mean thirty cases in a hundred) in both directions, and here the same structure appears under regulation with the trade-off made explicit and safety-critical. The addition pharma makes is that choosing the error well does not discharge the governance — it lowers the visible risk and leaves a residual that the qualification and the human backstop must cover, and both of those are things over-confidence in a well-tuned model can quietly let lapse.

The Lab network models only the deploying organization: its inspection model, its human inspectors as the backstop, and its inspection and batch records. No patient-safety or defect-escape outcome is computed on any diagram. The patients who receive the products are boundary-only; the asymmetric trade-off, the over-reject tuning, the emerging regulatory qualification, and the over-trust risk to the human backstop are institutional signals that live in this case file, never on any network. Consistent with the domain's caveat, nothing here is a claim that a named manufacturer's AI let a defect ship; the case is a portrait of a well-chosen trade-off and the two surfaces that keep its residual risk covered. The map's instruction is to credit the deliberate safe-direction tuning, to treat the AI-specific qualification as a real and unfinished check, and to protect the human backstop from the over-trust that would hollow it out at the one point that matters most.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

veillon2023aGroundingPeer-reviewedSave

Veillon, R., Shabushnig, J., Aabye-Hansen, L., et al. (2023). Applying Machine Learning to the Visual Inspection of Filled Injectable Drug Products. PDA Journal of Pharmaceutical Science and Technology, 77(5), 376-401. https://doi.org/10.5731/pdajpst.2022.012796 https://journal.pda.org/content/77/5/376

doi.org/10.5731/pdajpst.2022.012796

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance)

usfda2023GroundingGovernmentSave

U.S. FDA, CDER/OPQ (2023). Discussion Paper: Artificial Intelligence in Drug Manufacturing. Docket FDA-2023-N-0487. https://www.fda.gov/media/165743/download

https://www.fda.gov/media/165743/download

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: pharma_avi_inspection

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalMachine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic…

Machine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects that manual inspection or fixed-rule cameras would otherwise judge. In this safety-critical, regulated manufacturing setting the error trade-off is asymmetric and deliberate: a false accept — a missed defect in an injectable that reaches a patient — is a patient-safety failure, while a false reject — scrapping a good vial — is a cost, so the system is tuned to over-reject rather than risk a miss. Because the inspection sits inside a validated pharmaceutical quality process, the AI cannot simply be switched on; it must be qualified within that process, and a regulator is actively developing the framework for how AI in drug manufacturing should be validated and monitored.

veillon2023aGroundingPeer-reviewedSave

Veillon, R., Shabushnig, J., Aabye-Hansen, L., et al. (2023). Applying Machine Learning to the Visual Inspection of Filled Injectable Drug Products. PDA Journal of Pharmaceutical Science and Technology, 77(5), 376-401. https://doi.org/10.5731/pdajpst.2022.012796 https://journal.pda.org/content/77/5/376

doi.org/10.5731/pdajpst.2022.012796

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance)

usfda2023GroundingGovernmentSave

U.S. FDA, CDER/OPQ (2023). Discussion Paper: Artificial Intelligence in Drug Manufacturing. Docket FDA-2023-N-0487. https://www.fda.gov/media/165743/download

https://www.fda.gov/media/165743/download

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: pharma_avi_inspection

EmpiricalTwo governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a…

Two governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a validated quality process is not simply deployed but must be qualified and monitored for drift, and because the regulator's AI-specific framework is still developing, the qualification of the model's behavior over time is an emerging, not-yet-settled check rather than a solved one. Second, the human backstop: the manual inspector is what catches the false accepts the over-reject tuning is meant to avoid, so if inspectors come to defer to the AI and stop scrutinizing, that backstop erodes exactly where it matters most — the missed defect the asymmetric tuning was designed to prevent. The governable reading is that the over-reject tuning lowers the visible risk without removing it, and the qualification and the human backstop are what keep the residual risk covered.

veillon2023aGroundingPeer-reviewedSave

Veillon, R., Shabushnig, J., Aabye-Hansen, L., et al. (2023). Applying Machine Learning to the Visual Inspection of Filled Injectable Drug Products. PDA Journal of Pharmaceutical Science and Technology, 77(5), 376-401. https://doi.org/10.5731/pdajpst.2022.012796 https://journal.pda.org/content/77/5/376

doi.org/10.5731/pdajpst.2022.012796

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance)

usfda2023GroundingGovernmentSave

U.S. FDA, CDER/OPQ (2023). Discussion Paper: Artificial Intelligence in Drug Manufacturing. Docket FDA-2023-N-0487. https://www.fda.gov/media/165743/download

https://www.fda.gov/media/165743/download

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: pharma_avi_inspection