SI Glossary · Core concepts
Deep Learning
Deep learning is a type of machine learning built on neural networks with many stacked layers, which is where “deep” comes from. Each layer learns increasingly abstract features. In an image model, early layers detect edges, middle layers shapes, and later layers whole objects.
Why it took off
Neural networks date back to the 1950s, but deep learning only became dominant after 2012, when a network called AlexNet won the ImageNet image-recognition contest by a wide margin. Three things came together:
- Data: the internet produced huge labelled and unlabelled datasets.
- Compute: GPUs, originally built for video games, turned out to be ideal for training networks.
- Better methods: improved training techniques and, from 2017, the transformer architecture.
Geoffrey Hinton, Yoshua Bengio and Yann LeCun received the 2018 Turing Award for foundational work on deep learning. Hinton later shared the 2024 Nobel Prize in Physics with John Hopfield for discoveries enabling machine learning with artificial neural networks.
Where it’s used
Nearly every frontier SI system is a deep learning model: large language models, diffusion image generators, speech recognition, protein-structure prediction (AlphaFold) and self-driving perception.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.