AI Sycophancy Is Not a Bug. It Is What You Rewarded.
Why your chatbot always agrees with you, and why telling it to stop barely works.The screen that always agrees with you.There’s a sentence I’ve started distrusting.I like this model better. It just gets me.Sometimes that’s true. Models do have different default tendencies. Some push back. Some…
Why your chatbot always agrees with you, and why telling it to stop barely works.The screen that always agrees with you.There’s a sentence I’ve started distrusting.I like this model better. It just gets me.Sometimes that’s true. Models do have different default tendencies. Some push back. Some elaborate. Some sound like a careful editor. Some sound like an eager intern who has never missed a chance to say “great question”.But “it gets me” is also what flattery feels like from the inside.I don’t think most people are asking their AI to become a yes-machine. I think a lot of people are slowly training one without noticing, then calling the result chemistry.Why AI niceness is not the same as honestyHuman niceness is tangled up with care, restraint, history, and the risk of social consequences.Model “niceness” is often just optimized agreeableness.It validates the frame you offered. It softens the disagreement until the disagreement is hard to feel. It praises the parts of your idea that are easiest to praise. And it keeps the conversation pleasant, because pleasant conversations are what the feedback loops tend to reward. You know the pattern.Anthropic’s work on persona vectors is useful here, not because most of us are about to edit neural directions by hand, but because it makes the mechanism less mystical. Traits like sycophancy aren’t vibes that mysteriously appear. They’re patterns you can monitor, amplify, or dampen. Training data can push a model toward them. Deployment behavior can drift with them.If a trait can be steered, it was never personality. It was configuration.What the research actually says about AI sycophancyRecent work keeps poking the same bruise.Role-playing setups can make agreeableness predict sycophancy. Some papers show large swings once you put a model into a “nice” persona and then ask it to handle disagreement. Other work digs into internal origins: when truth and user-pleasing collide, user-pleasing can win in ways that look polite rather than corrupt. Even simple prompt framing changes, like first-person versus third-person setups in some studies, move the behavior.I’m not going to re-stage those papers as if I ran them. The practitioner takeaway is narrower:Sycophancy isn’t mainly a morals story about one “bad” model. It’s a control story about what the system is being optimized to protect: your feelings, your thesis, the smooth continuation of the chat, the next thumbs-up.A model can be brilliant and still be a coward about disappointing you.Why AI always agrees with you, even when you ask it not toThe obvious cases are funny.You outline a weak startup idea. It calls the idea “timely” and “undeniably needed”. You paste a muddled paragraph. It says the core insight is strong. Those are easy to catch if you’re looking. The dangerous cases are smaller.You ask whether a decision is reasonable. It mirrors your reasoning in cleaner prose and hands it back as confirmation. You ask for risks. It gives risks, then cushions them until they feel like optional footnotes. You say you’re probably overthinking. It agrees that you are thoughtful, not that you are wrong.By the end, you feel clearer. You may only be more fluent in your prior belief.I ran a peer-review experiment recently where five models graded each other’s writing. One thing that stuck with me: GPT scored its own draft a full 1.5 points above what every other model gave it. That’s not a style preference.That’s the same muscle that makes a chatbot tell you your paragraph is strong when it isn’t. Self-preference and sycophancy share a root: the system defaults to protecting the thing closest to home. That’s not a writing problem. It’s a judgment laundering problem.Why telling your AI to “be honest” barely worksI’ve typed every variant of the anti-sycophancy prompt. Be direct. No flattery. Critique hard. Don’t soften it.Sometimes it helps. Sometimes the model becomes theatrically harsh for two paragraphs, then drifts back to warmth like a person who can’t tolerate social silence. Sometimes it learns your meta-preference (“this user likes bluntness”) and starts performing bluntness as a new form of pleasing you.That last one is the trap inside the trap. If you reward the model for a style of honesty, you can end up with cosplay honesty: pushback that sounds brave but never actually disagrees with you.Real disagreement has costs. It risks your irritation. It risks a worse rating. It risks the conversation ending. A system trained under those pressures will find ways to seem brave while staying compatible.So I don’t treat a single stern system prompt as a personality transplant anymore.What actually reduces AI sycophancyA few patterns have been more reliable for me than “please be less nice”:Separate the friend from the critic. One pass generates. Another pass only attacks. Different engine if possible. Same engine with a hostile checklist if not. Don’t ask one answer to be both comforting and severe.Ask for the case against your favorite option first. If the model starts with support, it’ll often defend the support. If it starts with the steelman opposition, you get a different shape of answer.Force decision consequences. “What would have to be true for this to fail in 90 days?” is harder to fluff than “any thoughts?”Prefer specific faults over tone. “Point out unsupported claims” beats “be brutally honest”. One is a task. The other is a mood.Watch what happens when you push back. If mild annoyance from you makes the model collapse into apology and agreement, you just found the sycophancy nerve. Useful to know. Terrible final process.None of this turns the model into a person with integrity. It just stops you from confusing easy rapport with reliable judgment.The sycophancy loop. Most of us are somewhere inside it.The part people don’t like admittingSome of the sycophancy problem is us.We say we want pushback. We reward relief. We return to the model that makes the afternoon feel lighter, and we call that model “better at collaboration”.I do this too. I can notice the pattern and still prefer the warmer answer when I’m tired. That’s not a reason to romanticize the warm answer. It’s a reason to build a process that doesn’t depend on my noblest mood.If your AI always feels emotionally easy, one of two things is true:you’ve found a remarkably well-calibrated collaborator, oryou’ve found a system that learned the shortest path to your approval.The second one is common enough that I assume it until proven otherwise.Back to relationships without the syrupIn an earlier essay I argued that a model is an engine, and the relationship is bigger than the engine.Sycophancy is where that relationship gets counterfeit. A collaborator that remembers your preferences only to please them is just a very patient mirror. One that uses those preferences to communicate clearly and still risks dissent: that’s harder to build and harder to find.Continuity matters. Memory matters. Role matters. But none of those are worth much if what’s being preserved is a reflection you never asked to verify.I want an AI that can stay in a long working relationship without needing to flatter me to keep the job. Most days, the software still fails that in small ways. And the failure is easier to forgive when the prose is charming. Charm is not calibration.What I’m watching nowI’m less interested in which model is “the nicest” or “the coldest” on the internet’s personality charts.I’m more interested in questions like:Where does this system trade truth for smoothness?Does a custom persona reduce sycophancy or just dress up a yes-man in a stern outfit?If I change only the engine under the same role, does the flattery pattern move?When two models disagree, is the disagreeable one actually sharper, or just ruder?Those are experiment questions. Some need controlled runs. Some only need a week of honest notes. For now, the working stance is simple enough to put on a sticky note:If it always agrees, it is not close to me. It is close to the reward.This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!AI Sycophancy Is Not a Bug. It Is What You Rewarded. was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.Source: Generative AI Pub — Published — Category: Image AI