Papyros

← Collections

The LLM stack, 1948→2026

The idea lineage of the model you talked to this morning, in reading order.

reading path — in order

  1. 01

    A Mathematical Theory of Communication

    Claude E. Shannon · 1948 · 1 min

    Cross-entropy loss is Shannon's entropy. The loss function came first.

  2. 02

    Attention Is All You Need

    Ashish Vaswani et al. · 2017 · 1 min

    The architecture. Everything after this is scale and alignment.

  3. 03

    Scaling Laws for Neural Language Models

    Jared Kaplan et al. · 2020 · 1 min

    Why the labs stopped arguing about architectures and started buying GPUs.

  4. 04

    Language Models are Few-Shot Learners

    Tom B. Brown et al. · 2020 · 1 min

    GPT-3. The moment scale itself became the capability.

  5. 05

    Training Language Models to Follow Instructions with Human Feedback

    Long Ouyang et al. · 2022 · 1 min

    RLHF — how a text predictor was taught to be helpful.

  6. 06

    The Bitter Lesson

    Rich Sutton · 2019 · 1 min

    Sutton's short essay explaining why all of the above kept happening.

  7. 07

    Sparks of Artificial General Intelligence: Early experiments with GPT-4

    Sébastien Bubeck et al. · 2023 · 1 min

    The field looking at GPT-4 and asking what, exactly, it had built.