SI Glossary · Models & architecture
Multimodal Model
A multimodal model handles several “modalities” (types of data) at once. You can show it a photo of a broken appliance and ask what’s wrong, give it a chart and get a summary, or speak to it and hear a reply.
Examples of what multimodal models do
- Image understanding: reading screenshots, documents, handwriting and diagrams
- Audio: real-time voice conversation, transcription and translation
- Video: summarizing or answering questions about a clip
- Generation: producing images, speech or video alongside text
Why it matters
Early large language models were text-only. By 2026 nearly every frontier model accepts images as input, and many handle audio and video natively. Multimodality is essential for computer use: an agent operating a computer has to see the screen.
How they work
Most multimodal models convert each kind of input into tokens that live in a shared representation, then process them with a single transformer. Image and video generation usually still relies on diffusion models or similar techniques connected to the language model.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.