Skip to content
StrataHub

Algorithms · 1991

Mixture of Experts

In 1991 researchers proposed a network that divides labor among specialists, an idea now hiding inside the biggest AI models.

In 1991, Robert Jacobs, Michael Jordan, Steven Nowlan, and Geoffrey Hinton published 'Adaptive Mixtures of Local Experts', proposing that instead of one network learning everything, several smaller expert networks each specialize in part of the problem.

A small extra network, the gate, learns to route each input to the expert or experts best suited to handle it. The experts and the gate are trained together, so the division of labor emerges rather than being assigned by hand.

The appeal is efficiency and specialization. Different experts can capture different regimes of the data, and, crucially, not every expert has to run on every input. Only the chosen few are activated at a time.

That last property is why the idea roared back decades later. Modern language models can hold enormous numbers of parameters as a large set of experts, yet keep computation manageable by activating only a small slice per token, a design called sparse mixture of experts.

Systems like Google's Switch Transformer and the DeepSeek and Mixtral model families use exactly this trick to scale capacity far beyond what a dense network of similar running cost could reach.

It is a satisfying arc: an early-1990s idea about teams of specialists, largely dormant for years, now sits at the heart of some of the largest models ever built.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.