Ask a risk team why their new model underperformed in production and you'll hear about algorithms. Look at the postmortem and you'll almost always find data: a bureau attribute that changed meaning, a default flag defined three different ways across systems, a feature computed from data that wasn't actually available at decision time.
We've rebuilt enough credit risk pipelines to say it plainly: past a baseline of competence, model choice is worth a point or two of Gini. Data quality is worth ten. Here's where the bodies are buried.
Point-in-time correctness, or leakage by default
The single most expensive error in credit modeling is training on data as it exists now rather than as it existed at decision time. Bureau files get updated, delinquencies get cured and re-aged, balances change, disputes get resolved. If your training set joins today's warehouse tables to two-year-old applications, you are leaking the future into your features.
The symptom is a model that validates beautifully and disappoints immediately. We've seen offline-versus-production Gini gaps of 8–12 points traced entirely to point-in-time violations.
The fix is architectural, not statistical: snapshot-based feature stores where every feature value is stamped with its as-of date, and a training-set builder that can only join on as-of semantics. Make leakage structurally impossible rather than procedurally discouraged.
The default definition problem
"Default" sounds like a fact. It's a policy choice — 90+ DPD? Charge-off? Which products, what cure treatment, what performance window? In most institutions, three teams hold three incompatible definitions, and the model was trained on a fourth.
Before any modeling, we force a written target definition, signed off by risk and reconciled against finance's loss numbers. If the model's default rate doesn't tie back to the loss provision within explainable tolerance, stop. Nothing downstream matters until it does.
Bureau data is a supply chain, not a given
Bureau attributes arrive from external vendors with schema changes, coverage gaps, and semantic drift. Attribute definitions get revised. Thin-file and no-hit rates shift with your marketing mix. A bureau's data furnishers change what they report.
Treat it like the supply chain it is:
- Data contracts on every feed. Schema, ranges, null-rate expectations, and distribution checks enforced at ingestion — not discovered at retraining time.
- Vendor change monitoring. When the no-hit rate on a bureau pull moves 2 points week over week, that's a page to the data team, not a footnote in next quarter's model review.
- Reconciliation across sources. Where two sources should agree (bureau balance versus servicing system balance), measure the disagreement rate continuously. Rising disagreement is an early warning that one side is drifting.
Under SR 11-7, data quality is explicitly in scope for model risk management. Examiners increasingly ask not "is your model validated?" but "show us the controls on the data feeding it." Automated, logged data-quality checks are audit evidence, not just engineering hygiene.
Missingness carries signal — and risk
In credit data, a missing value is rarely random. A missing income field means something different on a digital application than on a broker-submitted one. A no-hit at the bureau is itself predictive. Blind median imputation destroys that signal and, worse, hides upstream breakage: when a source system starts silently dropping a field, imputation makes the pipeline look healthy while the model quietly degrades.
Our rule: missingness is modeled explicitly (indicator features, learned handling in tree models) and monitored explicitly (per-feature null-rate dashboards with alerting). The model should know a value is missing; the team should know why.
Reject inference and the population you can't see
Your labeled data only covers approved applicants. The model will score everyone. That through-the-door versus booked population gap widens every time credit policy or marketing changes, and it's a data problem before it's a methodology problem: most institutions don't retain declined-application data with anywhere near the rigor of booked accounts. Fix the retention and lineage first; argue about reject-inference techniques second.
What the discipline looks like in production
The credit data platforms we build share a spine:
- Versioned, as-of-correct feature store shared by training and serving — one computation path, no skew.
- Data contracts and automated quality gates on every ingestion, with failures blocking downstream jobs loudly.
- Full lineage from every production decision back through features to raw source records — which doubles as your adverse-action audit trail.
- Population and stability monitoring — PSI/CSI on inputs and scores, cut by channel and product, reviewed on the same cadence as model performance.
- A living eval harness so every data or model change is scored against frozen benchmark cohorts before release.
None of this is glamorous. All of it is the difference between a model that holds its Gini for three years and one that's quietly ten points worse than its validation report by month six.
Our typical entry point is a 4–6 week Pilot: a data-quality and leakage audit on one production risk model, with quantified findings — where the skew is, what it costs in discrimination power, and a prioritized remediation map. Co-Build stands up the feature store, contracts, and monitoring over 3–6 months. The model gets better as a side effect. It usually does.