Featured Concept · Reinforcement Learning

World
Models

The AI that dreams before it moves — 16 minutes of playable machines, zero equations.

Language Models 38 playable

Language Models · Free
Large Language Models
A machine that guesses the next word · 16 min
Language Models
Attention: Queries, Keys & Values
Every word interviews every other word · 15 min
Language Models
Self-Attention vs Cross-Attention
Same math, two sockets: itself or another sequence · 14 min
Language Models
Multi-Head Attention
Four specialist heads, one sentence — prune the dead ones · 14 min
Language Models
Induction Heads
The copy circuit behind in-context learning · 13 min
Language Models
Embeddings
Meaning becomes geometry — click the word map · 13 min
Language Models
Retrieval-Augmented Generation
Give the model a library card — then poison it · 13 min
Language Models
Vector vs Keyword Search
Type a query, watch two search engines disagree · 13 min
Language Models
Lost in the Middle
Why models forget the middle of long prompts · 12 min
Language Models
BPE Tokenization
Why AI can't spell strawberry or compare 9.11 vs 9.9 · 14 min
Language Models
Positional Encoding
Attention is order-blind — play with the fix · 14 min
Language Models
Rotary Position Embeddings (RoPE)
Position as rotation — spin Q and K yourself · 15 min
Language Models
The KV Cache
Why AI doesn't re-read your whole chat every word · 14 min
Language Models
Top-k & Top-p Sampling
Chop the probability tail before you roll the dice · 13 min
Language Models
Beam Search
Keep five guesses alive, not just one · 13 min
Language Models
Perplexity
How surprised is the model? Measure it · 12 min
Language Models
In-Context Learning
Teach a model with examples — no weights touched · 13 min
Language Models
Chain-of-Thought
Why writing the steps changes what it can solve · 14 min
Language Models
Base vs Instruct Models
Why raw GPT won't answer you — until it's taught to · 14 min
Language Models
LoRA Fine-Tuning
Retrain 0.1% of the weights, keep 99% of the gains · 13 min
Language Models
Knowledge Distillation
A tiny student learns the giant's every hunch · 14 min
Language Models
Quantization
Store weights in 4 bits and pray · 13 min
Language Models
Mixture of Experts
Huge model, tiny bill — route tokens to 2 of 8 experts · 14 min
Language Models
Speculative Decoding
A small model guesses, the big one checks — for free · 15 min
Language Models
Prompt Injection
When data gives orders — hijack an AI, then defend it · 14 min
Language Models
Hallucination Mechanics
Why AI makes things up — plausible is not true · 16 min
Language Models
Logits & Softmax
Raw scores to probabilities — the exponential megaphone · 14 min
Language Models
Temperature vs Top-p
One knob reshapes the odds, the other cuts them off · 12 min
Language Models
How BPE Is Learned
Watch a vocabulary grow, one merge at a time · 13 min
Language Models
Repetition Penalty
Tax words for showing up twice — break the doom loop · 13 min
Language Models
Text Watermarking
AI text signs its name in secret word choices · 14 min
Language Models
FlashAttention
Attention got 10x faster without changing the math · 15 min
Language Models
Attention Sinks
Half the attention lands on a meaningless first token · 14 min
Language Models
Test-Time Compute
What is 30 seconds of thinking worth? Now it has a price · 15 min
Language Models
Constrained Decoding
Force valid JSON by deleting illegal tokens, not asking nicely · 14 min
Language Models
Prompt Caching
Why the second identical prefix costs 10x less · 14 min
Language Models
Diffusion Language Models
A model that writes the whole page at once, then sharpens it · 15 min
Language Models
Mamba & State Space Models
A suitcase of memory instead of a library — attention's rival · 15 min

Generative Models 23 playable

Generative Models
Diffusion Models
Painting by un-destroying · 14 min
Generative Models
Latent Space
A whole face, compressed to two numbers you can drag · 12 min
Generative Models
VAE & the Reparameterization Trick
Backprop through a dice roll — with one line of algebra · 13 min
Generative Models
GANs: The Forger vs. the Detective
Two networks train each other — play both sides · 14 min
Generative Models
Mode Collapse
The forger learns one trick and makes only that · 13 min
Generative Models
Latent Interpolation
Morph two faces by walking the space between · 14 min
Generative Models
The U-Net Denoiser
Why diffusion's denoiser is shaped like a U · 13 min
Generative Models
Noise Schedules
How fast you add noise matters as much as the model · 13 min
Generative Models
DDPM vs DDIM Sampling
Drunk walk vs straight shot — same model, 10× faster · 14 min
Generative Models
Score Matching
Follow the arrows that point back to the data · 14 min
Generative Models
img2img & Strength
How much of your image survives the noise? · 12 min
Generative Models
Inpainting Masks
Repaint one region while the rest stays frozen · 13 min
Generative Models
ControlNet
Draw a skeleton the image must obey · 13 min
Generative Models
Flow Matching
Straighten the noise-to-image highway · 14 min
Generative Models
Autoregressive vs Diffusion
Pixel by pixel, or all at once? · 15 min
Generative Models
Video Temporal Consistency
Why AI video flickers, teleports and morphs over time · 14 min
Generative Models
Super-Resolution Hallucination
Upscalers don't recover detail — they invent it · 13 min
Generative Models
Style Transfer
Steal the brush, keep the scene · 14 min
Generative Models
Seeds & Determinism
Why the same prompt gives different images · 13 min
Generative Models
Negative Prompts
Steer the image away from what you don't want · 13 min
Generative Models
Memorization vs Generation
Did the model create that, or copy it? · 14 min
Generative Models
CFG Guidance Scale
The obedience dial — turn it too far and the image burns · 14 min
Generative Models
Multimodal Tokenizers
How a picture becomes 256 words a model can read · 14 min

Reinforcement Learning 21 playable

Reinforcement Learning
World Models
The AI that dreams before it moves · 16 min
Reinforcement Learning
Exploration vs. Exploitation
Play three slot machines with hidden odds, watch a live regret counter, then… · 14 min
Reinforcement Learning
Q-Learning
Watch a visible Q-table learn a gridworld: values flow backwards from the go… · 14 min
Reinforcement Learning
Reward Hacking
It maxes the score you wrote, not the goal you meant · 14 min
Reinforcement Learning
Credit Assignment
Reward lands at the end — which move earned it? · 14 min
Reinforcement Learning
Policy vs Value
Two ways to store the same intelligence: a value map that says how good ever… · 14 min
Reinforcement Learning
Monte Carlo Tree Search
Think by playing a thousand imaginary games · 14 min
Reinforcement Learning
Self-Play
Train a game-playing agent against copies of itself: watch skill emerge from… · 14 min
Reinforcement Learning
Discount Factor (γ)
How much is tomorrow worth to a machine? · 13 min
Reinforcement Learning
Sparse Rewards
Put the only reward at the exit of a maze and the agent wanders blind. Run i… · 14 min
Reinforcement Learning
The Sim2Real Gap
Train a hopping robot in a perfect simulator, deploy it onto messy real floo · 14 min
Reinforcement Learning
Model Predictive Control
Drive a toy car with Model Predictive Control: watch candidate futures fan o… · 14 min
Reinforcement Learning
Curriculum Learning
Watch a real learner fail at a hard task, then master it through staged diff… · 14 min
Reinforcement Learning
POMDPs & Belief States
Your agent sees a keyhole, not the world — act anyway · 15 min
Reinforcement Learning
ε-Greedy Exploration
One coin flip beats pure greed — until it doesn't · 13 min
Reinforcement Learning
Overfitting to the Simulator
It memorized this maze, not maze-solving · 14 min
Reinforcement Learning
Goal Misgeneralization
It learned the wrong goal — perfectly. · 14 min
Reinforcement Learning
Hierarchical RL
A manager picks subgoals; workers walk. · 13 min
Reinforcement Learning
Multi-Agent Emergence
Three dumb rules make a flock. Nobody wrote it. · 16 min
Reinforcement Learning
Tool Use & Agent Loops
Think, act, observe — run the loop, then break it · 14 min
Reinforcement Learning
Memory vs Reactive Agents
Without memory, every request is Groundhog Day. Teach an agent a fact and wa… · 14 min

Alignment & Safety 16 playable

Alignment & Safety
The RLHF Pipeline
Be the labeler — watch a model learn your taste · 16 min
Alignment & Safety
Reward Overoptimization: Goodhart's Law
Optimize the proxy too hard and quality collapses · 14 min
Alignment & Safety
Constitutional AI
A model that grades its own homework · 14 min
Alignment & Safety
Why Jailbreaks Work
Safety is a thin layer over capability · 14 min
Alignment & Safety
Sycophancy
RLHF taught it to tell you what you want to hear · 14 min
Alignment & Safety
Deceptive Alignment
An agent that behaves only while being watched · 15 min
Alignment & Safety
Interpretability Probes
Train linear probes on a toy network's activations, map where knowledge live · 14 min
Alignment & Safety
Feature Superposition
Pack five features into two neurons, slide the sparsity dial until they coll · 14 min
Alignment & Safety
Sparse Autoencoders
Unmix a model's thoughts into clean features · 14 min
Alignment & Safety
Refusal Boundaries
The fuzzy line between answer and decline · 13 min
Alignment & Safety
Prompt Injection vs Jailbreak
Jailbreak: the user attacks. Injection: the data attacks. · 13 min
Alignment & Safety
Scalable Oversight
How do you grade an exam you can't solve? · 16 min
Alignment & Safety
The Alignment Tax
Does safer mean dumber? Tune the dials and find out. · 13 min
Alignment & Safety
Specification Gaming
AIs obey the letter of your rules, never the spirit · 15 min
Alignment & Safety
Red-Teaming
Break a guarded model before the world does. · 14 min
Alignment & Safety
Process Reward Models
Grade every step, not just the answer — luck stops paying · 14 min

Training & Scaling 17 playable

Training & Scaling
The Loss Landscape
Where the ball lands was decided before it moved · 14 min
Training & Scaling
Backpropagation
The blame ledger: who pays for a wrong answer, and how much · 15 min
Training & Scaling
The Learning Rate
One dial: crawl, converge, or vaporize the run · 14 min
Training & Scaling
Overfitting vs Generalization
The crammer aces practice tests and fails the real one · 15 min
Training & Scaling
Dropout
Fire half your neurons at random — and get stronger · 14 min
Training & Scaling
Batch Size
Noisy but wise, or smooth but blunt — pick your poll · 14 min
Training & Scaling
Reading Loss Curves
Six diseases, one X-ray — play the diagnosis game · 15 min
Training & Scaling
Scaling Laws
The straight line that $100M training runs are bet on · 15 min
Training & Scaling
Emergent Abilities
The cliff that might be a staircase — settle it yourself · 15 min
Training & Scaling
Double Descent
Bigger gets worse, then weirder: error falls twice · 15 min
Training & Scaling
Grokking
Memorize in minutes, understand 4,000 steps later · 15 min
Training & Scaling
Curse of Dimensionality
In 1,000 dims, everything is far from everything · 14 min
Training & Scaling
Data Contamination
When the exam leaks into the textbook, scores lie · 14 min
Training & Scaling
Chinchilla Optimality
Bigger model or more data? The $10B allocation question · 15 min
Training & Scaling
GPU Parallelism
1.1TB model, 80GB card: clone it, slice it, pipeline it · 15 min
Training & Scaling
Model Merging
Average two models' weights, get one good at both · 14 min
Training & Scaling
Synthetic Data & Model Collapse
AI eating its own output: inbreeding, measured live · 14 min

Coming Soon in production

Training & Scaling
Scaling Laws
Coming soon
Training & Scaling
Gradient Descent
Coming soon
Training & Scaling
Backpropagation
Coming soon
Training & Scaling
Grokking
Coming soon
Training & Scaling
Double Descent
Coming soon
Training & Scaling
Chinchilla-Optimal
Coming soon