Datasets & Benchmarks · 2016
SQuAD
Stanford paid crowdworkers to write 100,000 questions about Wikipedia, then watched machines pass humans within two years.
In 2016, Stanford researchers Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang wanted a reading comprehension test that was both large and honestly answerable. Their solution, the Stanford Question Answering Dataset, hired crowdworkers to read passages from 536 Wikipedia articles and pose over 100,000 questions about them.
The design was elegant: every answer had to be an exact span of text from the passage. That made grading automatic and objective, no human judges needed, which in turn made a public leaderboard possible. The leaderboard became a phenomenon, with research groups worldwide racing for the top spot.
Progress was startling. Systems from Microsoft and Alibaba matched the human benchmark on exact-match accuracy in early 2018, less than two years after release. When BERT arrived later that year, it swept past every prior system, and SQuAD became the demonstration that pretrained transformers had changed the rules of natural language processing.
The authors responded to the saturation cleverly. SQuAD 2.0, released in 2018, added 50,000 unanswerable questions written to look answerable, forcing models to know when the passage does not contain the answer, a much harder and more honest skill.
SQuAD's deeper legacy is a caution the field keeps relearning: models beat the human baseline while still failing simple adversarial distractors, showing that leaderboard victory is not the same as understanding. Every modern language model evaluation, with its held-out answers and automatic scoring, owes something to SQuAD's template.
Related stories
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.