Skip to content
StrataHub

Manufacturing · January 19, 2026 · 6 min read

Predictive Maintenance: From Pilot to Production

Most predictive maintenance pilots die in the gap between a promising notebook and a plant floor that trusts the alerts. Here is how we close that gap.

Predictive maintenance is the most piloted and least productionized use case in industrial AI. Every plant we walk into has run at least one PdM experiment. Almost none of them have a model that a maintenance planner actually acts on every morning.

The gap is rarely the model. It is everything around the model: data plumbing, alert economics, and trust. We have taken PdM systems live on rotating equipment, presses, and CNC fleets, and the pattern of what kills pilots is remarkably consistent.

Why pilots stall

A typical pilot looks like this: a data scientist pulls six months of vibration and temperature data, trains an anomaly detector, and shows a slide where the model "would have caught" a bearing failure three weeks early. Leadership is impressed. Then nothing happens.

Three things went wrong, and all of them were invisible in the demo:

The historical data was hand-cleaned. Someone spent two weeks aligning timestamps, filling sensor gaps, and removing periods when the line was down. In production, that cleaning has to happen automatically, in near real time, every day. Nobody scoped that work.

"Would have caught" is not a metric. One retrospective catch tells you nothing about precision. When the model runs live and fires forty alerts a month, and thirty-five of them are noise, technicians stop reading them by week three. A PdM system that cries wolf is worse than no system, because it burns the credibility you need for the next attempt.

Nobody owned the response. An alert with no work-order path is a notification, not a maintenance strategy. If the alert does not land in the CMMS with a suggested action, a priority, and an owner, it evaporates.

What production-grade PdM actually requires

We scope PdM engagements around four layers, and we insist all four are in the plan before any model gets trained.

1. A boring, reliable data layer. Sensor data from historians (PI, Ignition, whatever the plant runs) needs automated quality checks: dead sensors, stuck values, clock drift, unit mismatches between assets. In one engagement, 11% of vibration channels were silently flatlined for weeks at a time. We built data-quality monitors before we built anything predictive, and they caught more real problems in month one than the model did.

2. Failure-mode framing, not generic anomaly detection. "Something is weird" alerts are cheap to build and expensive to operate. We work with reliability engineers to enumerate the failure modes that matter — bearing wear, misalignment, cavitation, tool wear — and build detectors per mode where labeled history exists, falling back to anomaly detection only for the long tail. A model that says "stage-2 bearing degradation, estimated 10–20 days to functional failure" gets acted on. A model that says "anomaly score 0.91" gets ignored.

3. Alert economics, tuned with the people who receive them. Before go-live we agree on an alert budget: how many alerts per week the maintenance team can realistically investigate. Then we tune thresholds to that budget and measure precision against it. On a recent deployment we held the line at roughly eight alerts per week across 120 monitored assets, with investigated-alert precision above 70% by month three. That number matters more than any offline AUC.

4. Closed-loop feedback. Every alert gets a disposition: confirmed, false positive, or inconclusive, captured in the work-order system. That feedback stream is the eval set for every future model improvement. Without it you are retraining blind.

If your PdM vendor cannot tell you the live precision of their alerts over the last 90 days, you do not have a production system. You have a pilot with a dashboard.

The 4–6 week pilot, done right

A pilot should de-risk production, not defer it. In our pilot engagements we deliberately do the unglamorous work first:

  • Week 1–2: connect to the historian, profile data quality across target assets, and pick 2–3 failure modes with usable history.
  • Week 3–4: build detectors, backtest against known failures and known-healthy periods, and report precision at a realistic alert budget — not just recall on the failures everyone remembers.
  • Week 5–6: run shadow-mode alerts into the actual maintenance workflow and measure whether planners can act on them.

The exit criterion is not "the model works." It is "we know the cost, the alert volume, the integration points, and the expected precision — and the maintenance team wants to keep it running." When a pilot ends with that sentence, the co-build phase that follows is engineering, not evangelism.

What the payoff looks like

When this works, the numbers are not subtle. On one rotating-equipment fleet, moving from calendar-based to condition-informed maintenance cut unplanned downtime on monitored assets by about a third over two quarters and eliminated a chunk of unnecessary preventive teardowns — each of which carried its own infant-mortality risk. The model was a modest gradient-boosted classifier over engineered spectral features. Nothing exotic. The value was in the plumbing, the framing, and the trust.

The takeaway

Predictive maintenance fails as a data science project and succeeds as an operations project with data science inside it. Budget most of your effort for data quality, alert economics, workflow integration, and feedback capture. Treat the model as the easy part — because relative to everything else, it is.

If your last PdM pilot ended with a slide deck instead of a work order, that is the diagnosis. The cure is scoping for production from day one.

Work with us

Shipping something like this?

We co-build production AI systems with enterprise teams — pilots in 4-6 weeks.