Skip to content
StrataHub

Algorithms · 2014

Attention

Attention was invented to help machines translate French novels, not to power chatbots.

In 2014, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio were working on machine translation. The systems of the day squeezed an entire source sentence into a single fixed-length vector, then tried to generate the translation from that summary. On long sentences it fell apart.

Their fix, published as neural machine translation that jointly learns to align and translate, let the model look back at the whole source sentence and, for each word it produced, decide which source words to focus on. That selective focus was named attention.

It solved a very human problem. To translate a French clause, you glance back at the relevant French words rather than trying to hold the entire sentence in your head at once. Attention gave the network the same freedom, and translation quality jumped, especially on long inputs.

Attention also softened the vanishing-gradient problem in sequence models by creating direct shortcuts between distant words, instead of forcing information through a long chain of steps.

Three years later, a Google team took the radical step of throwing away the recurrent network entirely and building a model out of attention alone. The result, the transformer, became the foundation of nearly every large language model that followed.

So the mechanism now driving chatbots, coding assistants, and image generators began as a modest trick for lining up words across two languages.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.