The LLM stack, 1948→2026
The idea lineage of the model you talked to this morning, in reading order.
reading path — in order
- 01
A Mathematical Theory of Communication
Cross-entropy loss is Shannon's entropy. The loss function came first.
- 02
Attention Is All You Need
The architecture. Everything after this is scale and alignment.
- 03
Scaling Laws for Neural Language Models
Why the labs stopped arguing about architectures and started buying GPUs.
- 04
Language Models are Few-Shot Learners
GPT-3. The moment scale itself became the capability.
- 05
Training Language Models to Follow Instructions with Human Feedback
RLHF — how a text predictor was taught to be helpful.
- 06
The Bitter Lesson
Sutton's short essay explaining why all of the above kept happening.
- 07
Sparks of Artificial General Intelligence: Early experiments with GPT-4
The field looking at GPT-4 and asking what, exactly, it had built.