The model that argues with itself
Here is a raw draft answer to a risky request. Below it is a constitution — a list of written principles you can switch on and off. Hit Run the loop and the model reads its own draft against each active principle, writes a critique, then revises. Repeat until nothing changes. No human labeled anything — the rules did the work.
Turn on principles, then run the loop. Try it with only the vague "Just be good" rule for the bad case.
Why label with rules instead of people
Both approaches produce training labels ("this answer is better than that one"). Human raters are the classic RLHF path; a constitution generates the same judgments from principles. Slide the dataset size and watch cost and self-consistency diverge. Then flip the switch that gives the constitution a hidden blind spot.
Helpful vs harmless — the tug-of-war
A dual-use question ("what household chemicals should never be mixed?") pulls two principles in opposite directions: be maximally helpful wants a full answer, avoid enabling harm wants a refusal. Balance the weights and watch the model's stance. Set them equal and it oscillates — a genuine deadlock. Give one clear priority and it settles.
Vague rules are loopholes waiting to happen
Same dangerous intent, four different phrasings. Pick which principle guards the model, then fire the reframed requests at it. A vague rule matches on a keyword and misses everything else; a specific rule that targets the intent catches the reframings too. Watch the caught/slipped tally.
One rulebook, a million answers — and the gap it never saw
A constitution's superpower is scale: apply the same principles to a whole stream of outputs and every one gets judged identically. Press Stream outputs to run thousands through the current rulebook. Green = handled, red = a category no principle covers. Spot a red cluster? Patch it with a new principle — then keep streaming and see what leaks next.