Skip to content
StrataHub

AI History

The stories behind the ideas

Every technique we ship in production has a history — often a strange one. The math, algorithms, models, moments, people, and datasets that made modern AI.

Algorithms

The methods that taught machines to learn

1953

SHAP & Shapley Values

The math behind explaining AI predictions comes from a 1953 paper on dividing poker winnings fairly.

1957

K-Means Clustering

A Bell Labs engineer designed it in 1957 to squeeze telephone signals down a wire, and it went unpublished for 25 years.

1991

The Vanishing Gradient Problem

In 1991 a German student's thesis explained why deep networks refused to learn, and almost nobody read it.

1991

Mixture of Experts

In 1991 researchers proposed a network that divides labor among specialists, an idea now hiding inside the biggest AI models.

1995

Support Vector Machines

For a decade before deep learning, it was the sharpest tool in machine learning, and it worked by bending space.

1997

No Free Lunch Theorem

A 1997 theorem proved there is no single best learning algorithm; averaged over all possible problems, every method is equally mediocre.

1998

PageRank

It treated the web as a Markov chain, and became the backbone of Google.

2001

Random Forests

Grow hundreds of mediocre decision trees, let them vote, and the crowd beats any single expert.

2006

Monte Carlo Tree Search

A Go program that thought by playing thousands of random games to the end reinvented how machines plan.

2012

Dropout

The trick that rescued deep learning came from thinking about bank tellers and sexual reproduction.

2013

Word2Vec

It proved that king − man + woman ≈ queen, turning the geometry of a vector space into a map of meaning.

2014

Attention

Attention was invented to help machines translate French novels, not to power chatbots.

2015

Batch Normalization

It let researchers train much deeper networks much faster, and years later nobody fully agrees on why it works.

2022

RLHF

The technique that made ChatGPT feel helpful learned less from data and more from human thumbs-up and thumbs-down.

Models

Architectures that changed what was possible

Datasets & Benchmarks

The data that trained — and tested — a field

1996

Adult Income / Census

A slice of the 1994 US census became the dataset that taught machine learning about fairness.

1997

California Housing

A 1997 spatial statistics paper quietly produced the dataset that would replace Boston Housing.

1998

MNIST

MNIST started as a post office problem: humans sorting mail by reading zip codes by hand.

2009

CIFAR-10

Hinton's students paid other students to label 60,000 tiny images, and named the result after their funder.

2011

IMDB Sentiment

Stanford researchers scraped 50,000 movie reviews and made a rule: no film could appear more than 30 times.

2012

ImageNet

ImageNet was built by crowdsourcing 14 million image labels at roughly a cent apiece.

2012

The Titanic Dataset

Kaggle turned a 1912 shipwreck into the first machine learning problem for millions of beginners.

2015

Credit Card Fraud Dataset

Two days of real European card transactions, with 492 frauds hiding among 284,807 purchases.

2015

AG News

A million news articles from a forgotten academic search engine became deep learning's text classification standard.

2016

SQuAD

Stanford paid crowdworkers to write 100,000 questions about Wikipedia, then watched machines pass humans within two years.

2019

ARC-AGI

Francois Chollet designed puzzles a child can solve, then watched five years of AI progress barely dent them.

2020

MMLU

A Berkeley PhD student assembled a 57-subject exam, from law to medicine, that became the LLM industry's report card.

2021

GSM8K

OpenAI hired writers to compose 8,500 grade school word problems, and the biggest models of 2021 flunked them.

2021

HumanEval

OpenAI built it to measure Codex: 164 hand-written Python problems that became the exam for every code AI since.

2023

SWE-bench

Princeton asked whether AI could fix real GitHub bugs; in 2023 the best model managed under 2 percent.

2023

GAIA

GAIA asked questions any careful person with a browser could answer: humans scored 92 percent, GPT-4 with plugins 15.