Datasets & Benchmarks · 2023
GAIA
GAIA asked questions any careful person with a browser could answer: humans scored 92 percent, GPT-4 with plugins 15.
In late 2023, researchers from Meta AI and Hugging Face, including Gregoire Mialon, Clementine Fourrier, Yann LeCun and Thomas Scialom, proposed an unusual test for artificial general intelligence. Instead of PhD-level puzzles, GAIA, short for General AI Assistants, asks 466 questions that are conceptually simple but practically tedious.
A typical GAIA question requires finding a figure in a specific report, cross-referencing it against a website, reading an attached spreadsheet or image, and combining the pieces over multiple steps. Each question has a single short factual answer, so grading is automatic and there is no partial credit for confident prose.
The founding result was the headline: human annotators scored 92 percent, while GPT-4 equipped with plugins managed about 15 percent. The authors argued this inversion, easy for humans and brutal for machines, was a more honest target than exams that models can pass through memorized knowledge.
GAIA is organized into three difficulty levels by how many steps and tools a question needs, which made it a natural obstacle course for the agent frameworks that proliferated from 2024 onward. Its leaderboard became a proving ground for systems that combine browsing, code execution and file handling.
Scores climbed steadily as agent architectures matured, and GAIA settled in alongside coding benchmarks as a standard measure of whether an AI assistant can actually do multi-step work rather than merely talk about it, the exact question enterprises ask before putting agents into production.
Related stories
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.