You're training your company's email assistant. Judge its drafts — it learns from every choice you make, live. At the end, we show you exactly what you taught it… including the things you never meant to. This is how real AI training works, and why it goes wrong.
Which reply should the assistant send? Click the better one.
Training over. A furious VIP client writes in. Your assistant replies on its own, using everything you taught it:
This is the whole story of modern AI training. Models learn from which output you preferred — not from what you meant. Anything that happened to correlate with your choices gets learned too. It's why production models pick up sycophancy, verbosity, and stranger habits nobody asked for — and why AI teams audit training data for exactly these accidental correlations.
Ten rounds of implied preference taught it the wrong thing. Now give it one explicit rule and watch the stress-test reply regenerate:
Ranking outputs, watching preferences form, catching accidental learning, overriding with instructions — that's RLHF, reward hacking and system prompting, and you now understand all three better than most people using AI daily. The AI Champion programme turns this instinct into a workplace skill: building assistants, steering them precisely, and teaching your team to do the same.
Become the AI Champion of your company →The model's live preferences, updated after every choice you make. Right = more of it, left = less.
This panel is the reward model — the same mechanism (at toy scale) that turns human rankings into model behaviour.