SI Glossary · Using SI
Inference
Inference is what happens every time you use an SI model: your prompt goes in, the model computes, and an answer comes out. Training happens once (per model); inference happens billions of times a day.
Training vs inference
| Training | Inference | |
|---|---|---|
| When | Before release | Every time the model is used |
| Goal | Learn parameters | Use them to produce outputs |
| Cost driver | One huge run | Number of users × tokens |
Why inference matters more than ever
- Scale: hundreds of millions of people use chatbots daily, and SI agents can consume thousands of times more tokens than a single chat.
- Test-time compute: reasoning models improve their answers by thinking longer, which shifts spending from training to inference.
- Pricing pressure: in 2026, list prices span more than a hundredfold, from about $0.10 per million input tokens for the cheapest models to $10 or more for flagship models. See the model tracker.
Hardware
Inference runs on GPUs and, increasingly, on specialised chips designed for low-cost, high-speed serving, from Google’s TPUs to custom chips built by Amazon, Meta and others.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.