AI visual inspection is the most demoed use case in manufacturing and one of the most abandoned. The pattern is depressingly consistent: a proof of concept hits 98% accuracy on a curated image set, the pilot goes to the line, and within eight weeks the operators have taped over the alert light. Not because the model was bad — because everything around the model was missing.
We've now shipped enough production quality-control systems to say plainly what separates the ones that stick from the ones that get taped over.
The demo-to-line gap
Three realities hit on day one of production that never show up in the notebook.
The line is not your dataset. Lighting shifts with the season and the skylight. A new material lot has a slightly different surface finish. Someone bumps the camera bracket during maintenance. Each of these moves the input distribution, and a model trained on three good weeks of images degrades silently.
Defects are rare and unevenly rare. A line running 2% scrap doesn't give you a balanced dataset; it gives you 50:1 imbalance overall and 1,000:1 for the defect classes that matter most — the ones that reach customers. The failure mode you most need to catch is the one you have twelve examples of.
False positives have a price tag with a short fuse. Every false reject is either scrapped good product or a manual re-inspection. If the system flags 5% of good parts on a line running 60 parts a minute, you have created a full-time re-inspection job and a floor that stops trusting the system in about two shifts. Operator trust, once spent, is very expensive to buy back.
What the production version looks like
Fix the physics before the model. Half of "hard" vision problems are lighting problems. Controlled, shrouded illumination, fixed working distance, and a trigger synchronized to part position remove more variance than any architecture change. We spend the first week of every inspection build on optics and fixturing, and it is the highest-ROI week of the project.
Anomaly detection first, classification second. With scarce defect examples, we start with models trained predominantly on good product — learning normality and flagging deviation. That catches novel defects a classifier has never seen, which matters because the defect that costs you a customer is usually the one that wasn't in the training set. Classification of known defect types layers on top as labeled examples accumulate, feeding scrap analytics and root-cause work.
Thresholds are a business decision, made explicit. Every deployment has a knob trading escapes against false rejects, and where it sits should be decided by quality and operations with the costs on the table — not defaulted by a data scientist. For a safety-critical characteristic, you bias hard toward false rejects and route flags to human review. For cosmetic defects on a B-surface, the calculus flips entirely.
The feedback loop is the system. Every human override — "that's not a defect, that's a water spot" — is a labeled training example. The production builds that keep improving have a one-tap disposition interface at the reject station, and those labels flow into a retraining pipeline with review gates. The builds that decay treat deployment as the finish line.
Evals and observability, factory edition
We hold quality-control models to the same production standard as any system we ship, with some domain specifics.
The eval suite is a versioned, growing library of real images — stratified by defect class, material lot, lighting condition, and line — with per-class recall requirements, not a single accuracy number. 99% overall accuracy is meaningless when the critical defect class has 60% recall. Every retrained model must clear the full suite, including a regression set of previously caught escapes, before promotion. Model versions are pinned and rollback is one command.
In production we monitor the model like a sensor: input drift (image statistics shifting means lighting or camera trouble before it means quality trouble), prediction-rate drift (a reject rate that halves overnight is a broken trigger, not a quality miracle), and override rates by shift and defect class. When overrides climb, either the process changed or the model drifted — both are findable within the day if the traces exist.
Where the industry is regulated — automotive, medical devices, aerospace — the system also has to fit the quality management system: documented validation against MSA-style criteria, change control on model versions, and audit trails linking every disposition to a model version and an image. We design for that from the start because retrofitting compliance onto an ML system is far more painful than building it in.
The economics, briefly
A well-scoped inspection cell — one defect family, one station — typically runs 4–6 weeks to a shadow-mode deployment on real hardware: that's a Pilot. The business case usually clears on some combination of escaped-defect cost (customer claims, sorts, recalls carry the big numbers), redeployed inspection labor, and the scrap-analytics byproduct — once every defect is detected, classified, and timestamped, root-cause work gets dramatically faster. One client traced a recurring seal defect to a specific upstream press within three weeks of go-live, using data the inspection system generated as a side effect.
Co-Build takes it across stations and defect families, on shared retraining and monitoring infrastructure, so each new cell costs less than the last. But the sequencing is non-negotiable: shadow mode, real confusion matrix, operator trust, then automation. A vision demo that hits 98% in the lab is a slide. A system the night shift still trusts in month six is a quality-control program.