The task board — drag the scale, watch abilities "switch on"
Eight tasks, one model family, one slider. Each cell shows measured accuracy (500 simulated test items per cell, all-or-nothing grading). Drag from small to large and you'll reproduce the famous experience: translation and summarization improve politely, while multi-step tasks — arithmetic, code, logic chains — sit dark for ages and then snap on within a single scale step.
Every task shares the same smooth underlying per-step skill; tasks differ only in how many steps must ALL be right. That single difference produces both curve shapes you see.
Three rulers, one truth — re-measure the same models
Take the scariest cliff from the board — 8-step arithmetic — and grade the exact same simulated models three ways: exact match (all 8 steps right), partial credit (fraction of steps right), and token likelihood (how much probability the model puts on the right answer). One click each. Same models, same task, three completely different stories.
This is the Schaeffer et al. (2023) argument in miniature: many published emergences turn linear when the metric gives partial credit. All three curves are measured from the same simulated model family, live.
Open the hood — cliffs are AND-gates over staircases
Why do multi-step tasks cliff even though nothing inside the model jumps? Because a composite task is an AND-gate: to answer a word problem you must parse it AND set up the equation AND do the arithmetic. Each subskill below improves smoothly at its own pace. The composite (gold) is their product — flat while any factor is weak, then steep once the last one catches up. Toggle the hood open and the "miracle" becomes arithmetic.
Composite success is computed as the product of the three subskill curves at every scale point. Speeding up subskill C moves the takeoff — the composite cliff is fully determined by its parts.
The counterpunch — some jumps are real phase changes
Now the other side of the debate, because it also wins sometimes. Here a toy network trains on a copy-pattern task, and inside it a circuit (think: an induction head) is slowly assembling. While the circuit is incomplete it contributes nothing — memorization does all the work. Press train: at some step the pieces click, the circuit completes, and accuracy on unseen patterns genuinely leaps — measured with a smooth ruler. No metric trick.
Conceptual simulation of grokking-style dynamics (circuit-formation with a threshold), not a real transformer — but the shape matches what interpretability work found for induction heads: a genuine internal phase change. More diverse data → the circuit pays off sooner → earlier transition.
The forecasting game — this fight has stakes
Why anyone cares: labs must predict at what scale a capability crosses a threshold — for product roadmaps and for safety commitments. Play forecaster: you see a capability's noisy history up to today (gray zone), and must click where it will cross 50%. Three rounds — but round types alternate between cliff-metric data and smooth-proxy data. Your average miss tells the story.
Each round generates a fresh capability curve (same AND-gate machinery as EP 03) with noise; your click is scored against the true crossing computed from the underlying curve.