SI Glossary · Models & architecture
Transformer
On this page
The transformer is the design that made today’s SI boom possible. It was introduced in the 2017 paper “Attention Is All You Need” by eight researchers at Google, and it is the “T” in GPT (Generative Pre-trained Transformer).
What made it different
Earlier language models (recurrent networks) read text one word at a time, in order, which was slow and made it hard to connect distant words. The transformer instead uses an attention mechanism to look at all the words in a passage at once and work out which ones matter to each other. For example, it can work out that “it” in “the trophy didn’t fit in the suitcase because it was too big” refers to the trophy.
Because the work happens in parallel, transformers train efficiently on GPUs. That made it practical to scale them to billions and then trillions of parameters.
Where transformers are used
- Large language models: GPT, Claude, Gemini, Llama, Grok, Qwen, DeepSeek
- Vision transformers for image recognition
- Multimodal models that mix text, images, audio and video
- Scientific models, including parts of AlphaFold for protein structure
Variants
Most frontier models in 2026 are decoder-only transformers, many using a mixture-of-experts layout. Researchers continue to explore alternatives, such as state-space models, that may scale better to very long inputs.
Sources
- Attention Is All You Need — Vaswani et al., Google, 2017
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.