Skip to content
SI.info

SI Glossary · Models & architecture

Transformer

Published 1 min read
On this page
  1. What made it different
  2. Where transformers are used
  3. Variants

The transformer is the design that made today’s SI boom possible. It was introduced in the 2017 paper “Attention Is All You Need” by eight researchers at Google, and it is the “T” in GPT (Generative Pre-trained Transformer).

What made it different

Earlier language models (recurrent networks) read text one word at a time, in order, which was slow and made it hard to connect distant words. The transformer instead uses an attention mechanism to look at all the words in a passage at once and work out which ones matter to each other. For example, it can work out that “it” in “the trophy didn’t fit in the suitcase because it was too big” refers to the trophy.

Because the work happens in parallel, transformers train efficiently on GPUs. That made it practical to scale them to billions and then trillions of parameters.

Where transformers are used

  • Large language models: GPT, Claude, Gemini, Llama, Grok, Qwen, DeepSeek
  • Vision transformers for image recognition
  • Multimodal models that mix text, images, audio and video
  • Scientific models, including parts of AlphaFold for protein structure

Variants

Most frontier models in 2026 are decoder-only transformers, many using a mixture-of-experts layout. Researchers continue to explore alternatives, such as state-space models, that may scale better to very long inputs.

← Back to the SI Glossary

Sources

  1. Attention Is All You Need — Vaswani et al., Google, 2017

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy