Datasets & Benchmarks · 2015
Credit Card Fraud Dataset
Two days of real European card transactions, with 492 frauds hiding among 284,807 purchases.
In 2015, researchers at the Universite Libre de Bruxelles, working with the payment processor Worldline, released something rare: genuine credit card transactions. The dataset covers two days of September 2013 purchases by European cardholders, 284,807 transactions in all, of which exactly 492 are fraudulent.
That ratio, 0.172 percent, is the whole point. A model that simply predicts 'not fraud' for everything is 99.8 percent accurate and completely useless. The dataset became the canonical demonstration that accuracy is the wrong metric for rare events, pushing learners toward precision, recall and area under the precision-recall curve.
Privacy shaped its unusual form. Because real financial details could not be published, the team led by Andrea Dal Pozzolo and Gianluca Bontempi transformed the original features with principal component analysis, releasing 28 anonymous numerical components plus only the transaction time and amount in the clear.
Hosted on Kaggle, it became one of the most downloaded datasets in the platform's history and the standard classroom for imbalanced learning: undersampling, oversampling, synthetic minority techniques like SMOTE, and cost-sensitive methods that weigh a missed fraud far more heavily than a false alarm.
The underlying research also tackled problems production teams still face, such as concept drift, since fraud patterns change as criminals adapt. For enterprises, the dataset endures as a compact reminder that in high-stakes detection problems, the events that matter most are almost always the rarest.
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.