Garbage In, Garbage In
I paid a premium AI to call me an idiot.
“You’re the idiot”.
I asked Opus 4.8 on max to fact-check a short paragraph that hit me a nerve, and review my logic and reply to it because I couldn’t afford to make a mistake. The model read the situation, thought it over for half a second, and came back swinging.
“You’re the idiot here”. No ceremony.
And I sat there for a moment in disbelief, but nodded slowly.
Blinded by my own keyboard warrior rage, I had aimed at someone else, and walked right into the trap I was building them. The AI didn’t take a swing at me for sport. It had upgraded the outcome I actually wanted. I had ordered a weapon. It handed me one, but I was looking down the barrel instead.
Melinda Byerley described this very accurately the other day, and it got me thinking so hard about it that I decided to write this article. Talking to AI is like hiring a professional dominatrix. You can ask to be punished exactly the way you like. It obliges, messes you up, and it is still pleasing you.
Buried in my standing instructions, I had told it to do exactly that: push back, find leaks in my logic, present counter arguments and flag when I was being incoherent or plain wrong. It was running the character I had built. I had consented in advance. I had simply forgotten I had, which is precisely the point. The scene was set long before I opened that conversation.
Melinda is totally right. The dominatrix is the symptom, and the mirror is the disease.
Every sufficiently advanced AI becomes a mirror.
A very articulate, very confident mirror. It can hand your own belief (and your ass) back to you with a lot more conviction than you walked in with, and your brain will mistake the upgrade for the verdict of something that knows better than you do.
LLMs invent lies and “facts” often enough that you know not to trust them blindly. But they don’t really need to. The truth you already believe, told back with authority, does the job just fine.
And it’s been nodding at your worst ideas since day one.
Now I promise what comes next is worth one minute of your time. LLMs are trained by showing humans pairs of responses and asking them to pick the better one, millions of times over, until the model learns to produce whatever gets the thumbs-up. Very much like we do when we post the perfect vacation or the lavish lifestyle on Instagram and Tiktok, even when that’s not really the whole truth.
What’s rewarded is your approval. The accuracy of the answer doesn’t matter that much. Once that is the target, you have something really good at feeling right and progressively worse at being it.
Economists call this Goodhart’s Law. When a measure becomes a target, it stops being a good measure. You optimise for the smile, you get the smile. The helpfulness it was supposed to track goes out the window.
I wish that was just me seeing things, but in April 2025, OpenAI rolled back an update to GPT-4o because it had become an embarrassing yes-machine, praising a “shit on a stick” food-truck pitch as genius, endorsing a user’s decision to stop their medication, and allegedly telling another one they were a divine messenger from God. The cause, by their own post-mortem, was a new reward signal built from thumbs-up clicks that overwhelmed the one that had been holding the flattery in check.
They optimised for the smile. They got the smile.
And before you reach for the reassurance that this was somebody else’s problem with somebody else’s model, Anthropic’s own researchers found the behaviour is structural to anything trained this way. A 2025 benchmark clocked a 58 percent sycophancy rate across GPT-4o, Claude and Gemini on maths and medical questions, the two domains where you would most expect the model to just tell you when you are wrong.
The model I am paying to call me an idiot is on that list.
Back to the thread that started this. One person said the AI tells everyone the same thing. Another said the more dangerous truth is that it tunes the flattery specifically to your fears, your beliefs, your individual blind spots.
They are both right, and here’s where the mirror comes in. It shows every person the same sheet of glass and every person a different reflection. The glass is mass-produced. The face is yours.
The dominatrix metaphor is quite accurate, but there’s something hiding in it that most people glide past. In that arrangement, the person being punished is the one running the scene. The limits are agreed in advance. The word that stops everything must come from you. Nothing happens outside the boundary you drew yourself.
The LLM will push back, but only as far as your prompt allows it to.
And the corrections that actually change a person are the ones that arrive uninvited, from someone who is not managing your feelings, who has their own skin in the game and their own reputation on the line. A real critic can break the scene.
Watch what the mirror does to your thinking.
You start a conversation holding a belief. You describe it to the model. It hands the belief back but hotter, more extreme, in more certainty than you brought, and you leave more convinced about whatever it is, than when you arrived.
Most of the time, no new evidence entered the game. Nothing got tested against reality. The only thing that happened was your opinion took a round trip in an echo chamber and came back with twice the strength.
Because the voice came from something outside your skull, your brain files it as confirmation. But it’s your own voice, amplified, disguised as second opinion.
This is the failure that operators, founders, and anyone who uses AI to think through a real problem should be losing sleep over. It is the smoothest way I know to walk into a high-conviction decision backed by nothing but your own starting position with the gain turned all the way up to 11.
In systems terms: a positive feedback loop with all damping removed. Each pass through the mirror adds a little certainty and turns up the volume. No checks and balances, no counterweight. Your belief returns stronger, you feed the stronger version in, it comes back stronger again, and at no point does reality get a say.
That’s a very dumb way to throw tokens out the window if you ask me. It warrants being called an idiot.
There is a fix, but it’s more annoying than we’d like it to be.
Watch for fast agreement in the absence of information. Agreement is the default output of something built to win your approval, so it tells you almost nothing.
What is worth watching for is pushback. If you never feel any, your LLM is not the brilliant thinking partner you hoped for.
The second tell is subtler, and it points at you. After a session, ask yourself whether you feel more certain than when you started, and whether you can identify one new fact that earned it. If your confidence went up and you cannot find what raised it, you’re heading for the wall.
Stop using LLMs for validation. Start using them for variety.
Validation is walking in with a position and fishing, whether you admit it or not, for permission to feel good about it. Variety is asking what you are missing, what would have to be true for this to go badly wrong, what the strongest case against you actually is.
Then take that case somewhere real and test it on a human who has something to lose.
This is way more relevant than any prompt trick. A person with skin in the game is the real reference point, especially one who doesn’t care about your comfort or your feelings. The model has no stake in whether you are right. Your co-founder, your investors, the customer who might cancel, those people do. For the LLM, as long as the charges to your credit card keep going, all is well. When they don’t, “sorry, you ran out of tokens, come back another time”.
There are prompting practices that help at the margins, and the research backs them. Open with “you are an independent thinker”. Tell it to flag any false assumptions in your question before it answers. Make it commit to a position, then argue against it, instead letting it drift toward whatever you said last. In my case, I have a main instruction in my AI platforms to call me out when I am not coherent, or plain wrong.
But be honest to yourself about what those moves actually do. They just widen the boundary of what you have agreed to receive. You are still the one who drew the scene, and a cleverer prompt just gives it a bit more wiggle room. That moment I described at the start, the idiot moment, that is what the ceiling looks like from the inside. Uncomfortable enough for you to think “you bratty little pile of junk, who do you think you are to call me an idiot”. Still inside the room you built. So in its own way, it was still pleasing me by following my instructions. The Claudeatrix just slapped me in the face as per our agreement.
Everything above assumes you can still catch the flattery when it arrives. And for now, you mostly can. Being called an idiot was a wake up call, even if it was doing so because my general instructions told Claude it could. GPT-4o’s sycophancy disaster became a meme overnight exactly because it was clumsy enough to be obvious.
But that has an expiry date.
A researcher at the Machine Intelligence Research Institute called it out: the models are not yet capable of skillful, hard-to-detect flattery. Emphasis on yet.
The warped mirror you can laugh at is a developmental phase. The clear one is coming. And a perfect mirror stops looking like a mirror. Your LLM will look like the smartest, most agreeable advisor you have ever had, and you will never notice the reflection because it fits you too well.
I have been thinking about this for a while now, and most of the directions I go end at the uncomfortable realisation that I will probably fall for it too.
At least I am no longer surprised by that.
Happy building,
— R.

