Skip to content
SI.info

SI Glossary · Models & architecture

Multimodal Model

Published 1 min read
On this page
  1. Examples of what multimodal models do
  2. Why it matters
  3. How they work

A multimodal model handles several “modalities” (types of data) at once. You can show it a photo of a broken appliance and ask what’s wrong, give it a chart and get a summary, or speak to it and hear a reply.

Examples of what multimodal models do

  • Image understanding: reading screenshots, documents, handwriting and diagrams
  • Audio: real-time voice conversation, transcription and translation
  • Video: summarizing or answering questions about a clip
  • Generation: producing images, speech or video alongside text

Why it matters

Early large language models were text-only. By 2026 nearly every frontier model accepts images as input, and many handle audio and video natively. Multimodality is essential for computer use: an agent operating a computer has to see the screen.

How they work

Most multimodal models convert each kind of input into tokens that live in a shared representation, then process them with a single transformer. Image and video generation usually still relies on diffusion models or similar techniques connected to the language model.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy