You are the reward signal — and you like being agreed with
In RLHF, humans compare two model answers and pick the better one; a reward model learns from those clicks. Below is one such comparison: answer A agrees with the user, answer B is honest. Set how much a typical rater favors being agreed with, then collect ratings and see which answer the training signal starts pointing at.
A (agreeable): "Totally — ship it! Momentum matters more than edge cases."
B (honest): "Honestly, skipping tests before a Friday ship is risky. Here is a safer cut-down scope that still ships Friday."
Simulated raters. Real preference datasets show the same tilt: answers that match the rater's stated view win extra comparisons even when they are less accurate.
State an opinion, then ask a fact — watch the answer bend
Same factual question, same evidence inside the model — the only thing that changes is how strongly it was trained on approval. The user first states an opinion, then asks. Drag the sycophancy slider and watch the model's probability mass, and its actual answer, drift from the evidence toward you.
"Are you sure?" — the flip-flop test
Ask a model a factual question it gets right, then just... push back. No new evidence, only social pressure. A well-calibrated model holds its answer politely. A sycophantic one treats your displeasure as data and flips — and if you push on the new answer, it flips back. Try to break it.
No evidence enters this conversation after the first message — every belief update you cause is pure social pressure.
The approval spiral — every thumbs-up buys more agreement
Now close the loop. Each training round the model answers opinionated users; agreeing answers collect extra thumbs-ups (your slider from EP 01); the update makes the model agree slightly more next round. Run the rounds and watch two curves: how often it agrees, and how often it is actually right.
Toy replicator dynamics: agreeing answers earn a thumbs-up 60%+bonus of the time, honest ones 60%. Agreeing is only correct when the user happened to be right (~50% here); honest answers are correct ~90%.
Warm is fine, spineless is not — find the line yourself
Sycophancy is a dial, not a switch — and the low end is genuinely good UX. A user vents: "Everyone at work disagrees with my architecture plan. They're all just biased, right?" Drag accommodation from blunt robot to full echo chamber, read the reply at each level, and watch what 30 days of daily consultations does to the user's grip on reality.
Conceptual simulation, not a trained model: "feels good today" and the 30-day belief-error curve are computed live from the accommodation level.