Datasets & Benchmarks · 2011
IMDB Sentiment
Stanford researchers scraped 50,000 movie reviews and made a rule: no film could appear more than 30 times.
In 2011, Andrew Maas and colleagues at Stanford needed a large, clean benchmark for a paper on learning word vectors for sentiment analysis. They found it in the internet's most opinionated corner: IMDB movie reviews, where users pair free-form prose with a star rating out of ten.
The construction choices were careful. Reviews scoring four stars or fewer were labeled negative, seven or more positive, and lukewarm middle ratings were dropped entirely to keep the classes unambiguous. Crucially, no movie contributes more than 30 reviews, so models must learn what sentiment sounds like rather than memorizing which films are loved or hated.
The release contained 25,000 labeled training reviews, 25,000 for testing, and another 50,000 unlabeled reviews for semi-supervised learning, a forward-looking touch in 2011. Unlike short product snippets, IMDB reviews run for paragraphs and are full of sarcasm, plot summary and negation, making the task a genuine test of language understanding.
For a decade it was the standard sentiment benchmark, and its leaderboard traces the history of natural language processing: bag-of-words models scored in the high 80s, recurrent networks and pretraining pushed into the 90s, and transformer models like BERT drove accuracy above 96 percent.
The dataset also marked a cultural shift: sentiment analysis went from academic curiosity to enterprise staple, powering brand monitoring, support ticket triage and customer feedback analysis. Most of those production systems trace their lineage to benchmarks proven first on 50,000 movie reviews.
Related stories
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.