SI Glossary · Models & architecture
Attention Mechanism
On this page
An attention mechanism is the core idea inside the transformer. For each word (or token) in a passage, attention computes how much every other token should influence its meaning, and blends information accordingly.
An intuition
In the sentence “The bank raised interest rates, so the river bank flooded with protesters”, the model needs to tell the two meanings of “bank” apart. Attention lets each “bank” look at nearby words (“interest rates” vs “river”) and update its representation based on the most relevant ones.
How it works, briefly
Each token produces three vectors: a query (“what am I looking for?”), a key (“what do I contain?”) and a value (“what do I pass on?”). Comparing each query with every key gives attention scores. Those scores decide how much of each value flows into the token’s new representation. Running many of these in parallel is called multi-head attention.
Why it matters
- It lets models capture long-range relationships in text, code and images.
- It runs in parallel, which makes training on GPUs efficient.
- Its cost grows with the square of the input length, which is why long context windows were hard. Much engineering since 2020 has gone into making attention cheaper for million-token inputs.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.