Skip to content
SI.info

SI Glossary · Using SI

Inference

Published 1 min read
On this page
  1. Training vs inference
  2. Why inference matters more than ever
  3. Hardware

Inference is what happens every time you use an SI model: your prompt goes in, the model computes, and an answer comes out. Training happens once (per model); inference happens billions of times a day.

Training vs inference

TrainingInference
WhenBefore releaseEvery time the model is used
GoalLearn parametersUse them to produce outputs
Cost driverOne huge runNumber of users × tokens

Why inference matters more than ever

  • Scale: hundreds of millions of people use chatbots daily, and SI agents can consume thousands of times more tokens than a single chat.
  • Test-time compute: reasoning models improve their answers by thinking longer, which shifts spending from training to inference.
  • Pricing pressure: in 2026, list prices span more than a hundredfold, from about $0.10 per million input tokens for the cheapest models to $10 or more for flagship models. See the model tracker.

Hardware

Inference runs on GPUs and, increasingly, on specialised chips designed for low-cost, high-speed serving, from Google’s TPUs to custom chips built by Amazon, Meta and others.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy