Skip to content
SI.info

SI Glossary · Using SI

Embeddings

Published 1 min read
On this page
  1. What embeddings are used for
  2. How they’re stored
  3. Inside models

An embedding turns something messy, like a sentence, a photo or a product, into a list of numbers (a vector) that captures its meaning. Items with similar meanings get similar vectors. “How do I reset my password?” and “I forgot my login” land near each other even though they share no words.

What embeddings are used for

  • Semantic search: find documents by meaning, not exact keywords
  • Retrieval-augmented generation: fetch the right context for a model
  • Recommendations: “customers who liked this also liked…”
  • Clustering and deduplication: group similar support tickets or remove near-duplicate data
  • Classification: route messages, detect spam or flag fraud

How they’re stored

Embeddings are kept in vector databases or search engines with vector support, which can quickly find the nearest neighbours among millions of items.

Inside models

Every large language model uses embeddings internally: each token is first converted into a vector, and the model’s layers transform those vectors as they process text.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy