AI Safety · 5 Interactives

The Bill for Being Careful

Every safety dial you turn on a model — refuse more, filter more, align harder — sends a bill to its capability score. Or does it? The five machines below let you pay the alignment tax yourself, find the frontier, and discover the strange point where more safety makes the world less safe.

Three dials The frontier The false tradeoff Shrinking the tax The backfire
EP 01

Three safety dials — and only one of them is expensive

You run a model lab. You have three ways to make your model safer: crank refusal strictness, train harder with RLHF alignment, or add an output filter. Each moves your model on the safety-vs-capability chart below. Move the sliders and watch the price tags — they are not the same. Cheap safety exists; you just have to buy it from the right dial.

Toy response surfaces, not measurements of any real model — but the shapes (saturating safety gains, convex capability costs, RLHF cheap until over-optimized) mirror what labs report.

The mental model: "the alignment tax" is not one number — it's a price list. Moderate RLHF is nearly free (it can even help usefulness), while brute-force refusal strictness gets exponentially expensive. This is why labs pour money into training-time alignment instead of just refusing more.
EP 02

The frontier — 140 possible models, and only a handful worth shipping

Here are 140 candidate models — every one a different random setting of the three dials from EP 01. Most are dominated: some other candidate beats them on safety and capability at once. Reveal the Pareto frontier, then drag the safety requirement and read the honest tax: the capability you truly must give up at each safety level.

Every dot's position is computed live from the same toy model as EP 01 — nothing is hand-placed.

Why it matters: arguments about "the" alignment tax usually compare a random dominated model to a random other one. The only fair question is the frontier's slope: at the safety level you need, what does the best-known recipe cost? Off-frontier configs aren't paying a tax — they're just badly tuned.
EP 03

The false tradeoff — a better technique beats you on both axes

The frontier is not a law of physics — it's just the best recipe you currently know. Slide the training effort along the naive recipe (keyword filters + blanket refusals), then unlock the better one (targeted alignment training). Watch a single green point land above and to the right of yours: safer and smarter. The tradeoff you were agonizing over was an artifact of a bad method.

Conceptual demo. Real-world versions: instruction-tuned models beating raw ones on helpfulness AND harmlessness; DPO/RLAIF recipes matching RLHF safety with less capability loss.

The lesson: when someone says "safety costs capability, period," ask which method. A technique can dominate another on both axes — the way instruction tuning made models more useful and less toxic than raw pretrained ones simultaneously. The tax isn't fixed; it's a research target.
EP 04

Watch the tax shrink toward zero — generation by generation

Fix a safety target, then advance through method generations — think keyword filters → RLHF → constitutional-style training → targeted fine-tunes. The gold line is the capability tax the best method of each generation pays at your target. The grey line is what you'd pay forever by just cranking refusal strictness. One of these gets cheaper every year. The other never does.

Stylized decay curve — real numbers vary by benchmark — but the direction matches reports: labs now regularly ship aligned models that match or beat their unaligned baselines.

Good case: for well-understood harms, modern alignment methods have pushed the measured tax to roughly zero — some aligned models score higher because alignment teaches instruction-following too. Bad case: the grey line — strictness-only "safety" — pays full price every generation and buys the least.
EP 05

The backfire — when a safer model makes the world less safe

Final twist: your model doesn't exist in a vacuum. It has 1,000 users, and an unmoderated alternative exists one download away. Crank refusal strictness and watch two numbers fight: harm caused by your model falls, but frustrated users defect to the unsafe alternative — and total harm across the ecosystem can rise. Find the strictness that actually minimizes harm.

Toy ecosystem model: benign-refusal rate, user migration and harm rates are invented — but the U-shape is real whenever users have an unsafe outside option.

The uncomfortable truth: an over-refusing assistant doesn't eliminate risky queries — it exports them to models with no safeguards at all. Past the optimum, every extra point of strictness increases total harm. Real safety is a systems property, not a slider you max out.
Keep playing
Alignment & Safety
Scalable Oversight
How do you grade an exam you can't solve?
Alignment & Safety
Specification Gaming
AIs obey the letter of your rules, never the spirit