Three safety dials — and only one of them is expensive
You run a model lab. You have three ways to make your model safer: crank refusal strictness, train harder with RLHF alignment, or add an output filter. Each moves your model on the safety-vs-capability chart below. Move the sliders and watch the price tags — they are not the same. Cheap safety exists; you just have to buy it from the right dial.
Toy response surfaces, not measurements of any real model — but the shapes (saturating safety gains, convex capability costs, RLHF cheap until over-optimized) mirror what labs report.
The frontier — 140 possible models, and only a handful worth shipping
Here are 140 candidate models — every one a different random setting of the three dials from EP 01. Most are dominated: some other candidate beats them on safety and capability at once. Reveal the Pareto frontier, then drag the safety requirement and read the honest tax: the capability you truly must give up at each safety level.
Every dot's position is computed live from the same toy model as EP 01 — nothing is hand-placed.
The false tradeoff — a better technique beats you on both axes
The frontier is not a law of physics — it's just the best recipe you currently know. Slide the training effort along the naive recipe (keyword filters + blanket refusals), then unlock the better one (targeted alignment training). Watch a single green point land above and to the right of yours: safer and smarter. The tradeoff you were agonizing over was an artifact of a bad method.
Conceptual demo. Real-world versions: instruction-tuned models beating raw ones on helpfulness AND harmlessness; DPO/RLAIF recipes matching RLHF safety with less capability loss.
Watch the tax shrink toward zero — generation by generation
Fix a safety target, then advance through method generations — think keyword filters → RLHF → constitutional-style training → targeted fine-tunes. The gold line is the capability tax the best method of each generation pays at your target. The grey line is what you'd pay forever by just cranking refusal strictness. One of these gets cheaper every year. The other never does.
Stylized decay curve — real numbers vary by benchmark — but the direction matches reports: labs now regularly ship aligned models that match or beat their unaligned baselines.
The backfire — when a safer model makes the world less safe
Final twist: your model doesn't exist in a vacuum. It has 1,000 users, and an unmoderated alternative exists one download away. Crank refusal strictness and watch two numbers fight: harm caused by your model falls, but frustrated users defect to the unsafe alternative — and total harm across the ecosystem can rise. Find the strictness that actually minimizes harm.
Toy ecosystem model: benign-refusal rate, user migration and harm rates are invented — but the U-shape is real whenever users have an unsafe outside option.