Every fraud team we meet has the same artifact: a rules engine with 800 to 3,000 rules, accreted over a decade, where nobody can safely delete anything. Rule 447 was written after a 2019 incident by an analyst who left in 2021. It fires 400 times a day. Nobody knows if it still catches anything.
The instinct is to rip it out and replace it with a model. That instinct is wrong, and acting on it is how fraud-ML programs fail. Here's the migration path we've seen work.
Rules aren't the enemy — unmeasured rules are
Rules have real virtues: they're instant to deploy when a new attack pattern appears, they're trivially explainable, and compliance-mandated controls (sanctions screening, velocity limits, geographic blocks) must be deterministic. No regulator wants to hear that your OFAC control is probabilistic.
The problem isn't rules. It's that rule estates are typically unmeasured. Before we build any model, we instrument the incumbent: per-rule precision, catch contribution, and overlap analysis over 6–12 months of labeled outcomes. The findings are always the same shape — in one engagement, 30% of rules had not caught a single confirmed fraud in a year while generating a third of all alerts. That's not a modeling insight. That's an observability insight, and it pays for itself before any ML ships.
What ML actually adds
Models earn their place on the patterns rules can't express: interactions across dozens of features, behavioral baselines per customer, subtle sequences of events. A gradient-boosted or deep model scoring on top of a feature store — device fingerprint history, merchant patterns, velocity aggregates over multiple windows, network features linking accounts — routinely finds fraud that no human-writable rule combination would.
The realistic production wins we target: 20–40% reduction in false positives at constant fraud catch, or equivalently a meaningful catch-rate lift at constant review capacity. False positives are the number to obsess over. Every false decline is a legitimate customer insulted at checkout, and every unnecessary alert burns analyst hours — alert fatigue is how real fraud slips through a fully staffed team.
The hybrid architecture
What we deploy is layered, not either/or:
- Deterministic layer. Compliance controls and high-confidence blocks stay as rules. Non-negotiable, versioned, audited.
- Model layer. A calibrated risk score from the ML model, computed within a hard latency budget — for card-present authorization that means the full path, feature retrieval included, in under 50–100ms at p99. Feature freshness and serving latency are engineering problems that kill more fraud models than AUC ever does.
- Decision layer. A policy engine combines score, rule outcomes, and business context into approve / step-up / review / decline. Thresholds live here, owned by the fraud team, changeable without a model release.
Training-serving skew is the classic silent failure in fraud ML. If features are computed one way in the offline training pipeline and another way in the real-time path, your offline metrics are fiction. Build the feature store once, use it for both, and test parity continuously.
Labels lie, and adversaries adapt
Two hard truths that shape everything downstream.
First, fraud labels arrive late and dirty. Chargebacks take 30–90 days to materialize; some fraud is never reported; blocked transactions have no outcome at all, which quietly biases retraining. We handle this with explicit label-maturity windows, careful treatment of declined traffic (including small controlled holdout approvals where risk appetite allows), and by never evaluating a model on labels younger than the maturity window.
Second, fraud is adversarial. The distribution shifts because intelligent opponents probe your defenses. That makes monitoring and retraining cadence part of the core design: drift detection on inputs and score distributions, weekly performance readouts against matured labels, and a champion/challenger harness so a retrained model proves itself in shadow before taking traffic. When a novel attack lands at 2 a.m., analysts ship a rule immediately — and that rule's catches become labeled training data for the next model version. The rules layer becomes the rapid-response system feeding the learning system.
The migration sequence
- Instrument the incumbent (weeks, not months). Per-rule metrics, alert outcomes, review-queue analytics. Kill the provably dead rules.
- Shadow-score. Deploy the model with zero decision authority. Compare against rules on live traffic for 4–8 weeks with matured labels.
- Take one decision. Give the model authority over a single segment — typically the gray zone where rules currently punt to manual review. Measure, expand.
- Converge on the hybrid. Rules for compliance and rapid response, models for pattern detection, one decision layer with full audit trails on every decision path.
At StrataHub, step one and a shadow-scoring model is a Pilot: 4–6 weeks, running against your live traffic, with a quantified comparison of ML versus your current rule estate. Co-Build takes it through the decision layer, the feature store, and the retraining loop over 3–6 months. Production-or-nothing applies double in fraud — a model that only works in a notebook is a model the fraudsters never meet.
Don't decommission the rules engine. Demote it to the jobs it's actually good at, put a measured model beside it, and let the labeled outcomes decide who owns each decision.