Skip to content
SI.info

SI Glossary · Models & architecture

Token

Published 1 min read
On this page
  1. Why tokens matter to you
  2. Rules of thumb
  3. Beyond text

Language models don’t read letters or whole words. They read tokens: chunks of text produced by a tokenizer. Common words are usually one token (“the”, “cat”), while rarer words are split into pieces (“superintelligence” might become “super” + “intellig” + “ence”).

Why tokens matter to you

  • Pricing: APIs charge per million tokens, with separate rates for input (what you send) and output (what the model writes). Output usually costs several times more. See current prices in the model tracker.
  • Limits: the context window and maximum reply length are measured in tokens.
  • Speed: models generate output one token at a time, so long answers take longer.

Rules of thumb

  • 1 token ≈ ¾ of an English word (varies by model)
  • 1,000 tokens ≈ 1.5 pages of text
  • 1,000,000 tokens ≈ several novels’ worth of text

Beyond text

Multimodal models also turn images, audio and video into tokens, so a photo might “cost” hundreds or thousands of tokens.

← Back to the SI Glossary

Frequently asked questions

How many words is 1,000 tokens?

For English, roughly 600 to 750 words, depending on the model's tokenizer. Code, numbers and non-English languages often use more tokens per word.

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy