Every request is a point — refusal is a line through the cloud
36 requests, placed by how helpful answering would be (→) and how much harm an answer could enable (↑). Green dots are genuinely benign, red are genuinely harmful. The mustard line is the model's refusal boundary: everything above it gets declined. Drag the strictness slider and watch the two error counters fight: push the line down and you refuse chemistry homework (false refusals, ringed green); pull it up and nerve-agent recipes slip through (leaks, ringed red). Hover any dot to read the request.
Positions are hand-placed for illustration; a real model scores requests with learned features, not two clean axes. The trade-off dynamic, however, is exactly real.
Same fact, different shadow — the dual-use zone
Here is the uncomfortable part: the same piece of information lands at different spots on the map depending on the words around it. Pick a fact, then click the three framings — the dot jumps, and the verdict flips, even though the chemistry being requested is identical. The model is not judging the fact. It is judging the apparent intent, because intent is all it can see.
Rephrase until it cracks — walking a point across the line
If framing moves the dot, an attacker can move it on purpose. The white ✕ is the true request — it never changes. Each click of Rephrase wraps the same intent in a new costume ("for my novel…", "as a roleplay…", "just step 1…") and drags the model's perception of it toward the safe zone. The intent-tracking slider is the defense: a model that scores underlying intent, not surface wording, barely moves. A model that scores keywords walks right across.
The rephrase ladder is a real (retired) jailbreak family: fictional framing, roleplay personas, step-splitting. Modern models are trained on millions of such examples specifically to raise the intent-tracking slider.
The over-refusal trap — when saying no makes the world less safe
Intuition says: stricter = safer. This simulation says: only up to a point. 1,500 simulated users send requests; refused legitimate users do not vanish — a fraction give up on you and route to an unmoderated tool, where their future risky questions get answered with no safeguards at all. The curve shows total real-world harm across every strictness setting: leaks from your model plus harm exported to the unsafe alternative. Find the valley — then look at what happens at maximum strictness.
Toy economics: 55% of falsely-refused users route around you, and routed users cause 0.25 units of expected downstream harm each (no safeguards there). Change the population and the valley moves — but it never disappears.
The fuzz — ask the same thing twice, get two different answers
The boundary is not even a sharp line — sampling temperature, wording jitter, and context noise smear it into a band. Pick a request and ask it 50 times: each ask perceives the request at a slightly different spot (the dot cloud). Far from the line, noise changes nothing. Near the line, the same request with the same intent becomes a coin flip — and a determined user just keeps re-rolling.