Papyros

2020

Language Models are Few-Shot Learners

GPT-3: the moment scale alone became capability, and prompting replaced fine-tuning.

“We train GPT-3, an autoregressive language model with 175 billion parameters, and test its performance in the few-shot setting.”

Scale as capability

Brown et al. trained GPT-3 at 175B parameters and found that many tasks improve smoothly with scale. The model does not need fine-tuning for every new benchmark. A short prompt with a few examples is often enough.

In-context learning

The surprising behavior is few-shot learning: the model reads a task description and a handful of input-output pairs in its context window, then completes the next item. No weight update occurs. The task is inferred at inference time.

What it changed

GPT-3 made scale feel like a product strategy. It also exposed the limits of prompt-only learning: strong on many benchmarks, brittle on others, and expensive to run. Every later frontier model inherits this paper's question: what does scale buy you, and what still requires training?

← previous · 2019

The Bitter Lesson

next · 2020 →

Scaling Laws for Neural Language Models