SI Glossary · Using SI
Embeddings
On this page
An embedding turns something messy, like a sentence, a photo or a product, into a list of numbers (a vector) that captures its meaning. Items with similar meanings get similar vectors. “How do I reset my password?” and “I forgot my login” land near each other even though they share no words.
What embeddings are used for
- Semantic search: find documents by meaning, not exact keywords
- Retrieval-augmented generation: fetch the right context for a model
- Recommendations: “customers who liked this also liked…”
- Clustering and deduplication: group similar support tickets or remove near-duplicate data
- Classification: route messages, detect spam or flag fraud
How they’re stored
Embeddings are kept in vector databases or search engines with vector support, which can quickly find the nearest neighbours among millions of items.
Inside models
Every large language model uses embeddings internally: each token is first converted into a vector, and the model’s layers transform those vectors as they process text.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.