Skip to content
StrataHub

Algorithms · 2013

Word2Vec

It proved that king − man + woman ≈ queen, turning the geometry of a vector space into a map of meaning.

In 2013, Tomas Mikolov and colleagues at Google released word2vec, a method that represents each word as a vector of a few hundred numbers, learned by predicting which words tend to appear near which others.

The startling result was that meaning turned into geometry. Take the vector for 'king', subtract 'man', add 'woman', and the nearest vector is 'queen'. Similar analogies held for capitals and countries, verb tenses, and comparatives. Relationships between words became consistent directions in space.

The underlying intuition is old and linguistic: you shall know a word by the company it keeps. Words used in similar contexts end up with similar vectors. Word2vec made this idea trainable at scale, on billions of words, using a deliberately shallow and fast model.

Those vectors, called embeddings, could be fed into other systems as a rich numerical starting point, giving machines a usable sense of which words are related before any task-specific training began.

Word2vec kicked off the embedding era. The same idea now underlies how modern language models represent not just words but sentences, images, and more, all as points in high-dimensional space where nearness means similarity.

It also carried a warning. Because the vectors absorb whatever patterns are in the text, they absorb its biases too, a finding that launched a whole line of research into fairness in learned representations.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.