Skip to content
StrataHub

Datasets & Benchmarks · 1997

California Housing

A 1997 spatial statistics paper quietly produced the dataset that would replace Boston Housing.

In 1997, statisticians Kelley Pace and Ronald Barry published a paper on sparse spatial autoregressions in a statistics journal, illustrating their method with data they had assembled from the 1990 US census. Few could have guessed their example dataset would outlive the paper by decades.

The data describes 20,640 California census block groups, each summarized by median income, housing age, average rooms and bedrooms, population, household occupancy, latitude and longitude. The target is the median house value of the block group, and one quirk is visible in every scatter plot ever made from it: values were capped at 500,001 dollars, producing a hard ceiling line.

The dataset spread widely after it was bundled into scikit-learn, the dominant Python machine learning library, as a ready-made regression example. Its role grew in 2021 when scikit-learn removed the older Boston Housing dataset over a feature that encoded a racially discriminatory assumption, and California Housing became the default regression demo for a generation of tutorials.

It teaches well because geography matters. House values depend on proximity to the coast and to cities in ways that a straight line through income cannot capture, so the dataset cleanly shows why tree ensembles and neural networks beat plain linear regression, and how latitude and longitude interact as features.

It is a small lesson in how infrastructure shapes science: a convenient example, packaged in the right library at the right time, quietly trained millions of people in the basics of predictive modeling.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.