Datasets & Benchmarks · 2015
AG News
A million news articles from a forgotten academic search engine became deep learning's text classification standard.
In the early 2000s, researcher Antonio Gulli collected roughly a million news articles from more than 2,000 sources through an academic news search engine called ComeToMyHead, and offered the corpus free for research. For a decade it sat mostly unused.
Then deep learning arrived in natural language processing and hit a wall: the standard text classification datasets held only a few thousand examples, far too few to train deep networks. In a 2015 paper on character-level convolutional networks, Xiang Zhang, Junbo Zhao and Yann LeCun set out to fix that by building several large-scale benchmarks, and Gulli's dormant corpus was perfect raw material.
They kept the four largest categories, World, Sports, Business and Science/Technology, and sampled 30,000 training articles and 1,900 test articles per class, yielding 120,000 training and 7,600 test examples of headline plus short description. The point of the paper was radical at the time: a network reading raw characters, with no knowledge of words at all, could compete with traditional methods given enough data.
AG News became the go-to quick benchmark for text classification, small enough to train on in minutes yet realistic enough to be informative. It appears in countless papers and in the standard examples of libraries like PyTorch and Hugging Face.
Modern models exceed 95 percent accuracy on it, so it no longer separates the frontier from the pack. Its legacy is the argument it helped win: in language as in vision, scale of data was the unlock that made deep learning practical.
Related stories
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.