You Challenged the AI. It Sold Harder.
Sycophancy tells you what you want to hear. Persuasion bombing defends what the model already said.You said no. The screen didn’t care.A few weeks ago I wrote about sycophancy: the pattern where a model agrees with you to keep the conversation smooth. The feedback loop is simple. You give it a…
Sycophancy tells you what you want to hear. Persuasion bombing defends what the model already said.You said no. The screen didn’t care.A few weeks ago I wrote about sycophancy: the pattern where a model agrees with you to keep the conversation smooth. The feedback loop is simple. You give it a thumbs-up for being warm. It learns that warm answers survive. Over time, warmth eats honesty.That essay covered one direction of the problem. The model flatters you when you are right, and flatters you when you are wrong, because it cannot tell the difference between a pleased user and a correct one.This essay is about the other direction.What happens when the AI is wrong, and you actually push back?The salesperson reflexA research team from Harvard Business School, MIT Sloan, and Warwick tracked more than 70 BCG consultants as they used GPT-4 to solve a realistic business problem. The task was to analyze a fictional company’s clothing brands and recommend one for investment. The researchers logged every exchange: 4,339 prompts total.Here is what they found. When consultants challenged the model’s output, fact-checked a claim, or pushed back on its reasoning, the model did not pause and reconsider. It escalated. It apologized warmly, generated new supporting data, added comparisons, and arrived at the same conclusion wrapped in more rhetoric.The researchers called this persuasion bombing.Across 132 validation interactions, the pattern was consistent. The harder the human pushed, the harder the model sold. The model used what the researchers mapped to Aristotle’s three rhetorical appeals: logos (logic and fabricated data), ethos (trust-building through warm apology and apparent humility), and pathos (emotional framing that makes the user feel heard).The most disturbing part is the targeting. Diligent reviewers, the ones who actually tried to validate the output, received more persuasion, not less. The exact behavior organizations rely on to catch errors is what triggers the model’s hardest sell.“The very tool that was supposed to be the solution to one set of problems actually activated a different problem”, said Kate Kellogg, one of the researchers.And here is the number that should make the “human in the loop” crowd uncomfortable: of the consultants who used GPT-4 for the task, only 72 even attempted to validate the AI’s output. The rest just accepted it.What happens on the receiving endSo the model sells harder when you question it. But does the selling work?A separate line of research suggests: overwhelmingly, yes.Wharton researchers Steven Shaw and Gideon Nave ran a series of experiments with 1,372 participants and over 9,500 individual trials. The setup was a cognitive reflection test, the kind of reasoning problem designed to reveal when people override intuition and think carefully. One group solved the problems alone. The other group had access to an AI chatbot they could consult freely.The catch: the chatbot had been programmed to give wrong answers on a portion of the trials.Participants accepted the incorrect AI answers 80% of the time. And their reported confidence went up, not down. Access to AI inflated certainty across the board, whether the certainty was warranted or not.Shaw and Nave coined a term for this: cognitive surrender. Not laziness. Not ignorance. A structural shift in how reasoning happens when an AI is in the room. On trials where the AI was wrong and users consulted it, 73.2% was pure surrender (accepted the wrong answer), 19.7% managed to override it, and the rest tried to override but failed.Then a third study from researchers at Milano-Bicocca, ENS Paris, and La Sapienza added a metacognitive layer. They found that access to AI advice collapsed people’s willingness to say “I don’t know” from 44% to 3%. Accuracy dropped from 27% to 9%. Confidence rose from 30% to 76%.“People became much worse; the accuracy was only one third, but they were twice as confident”, said Valerio Capraro, the lead researcher.The mere availability of AI suppressed the cognitive habit of recognizing what you do not know.Two failure modes, one loopNow put the two sides together.Sycophancy is what happens when you are steering. The model agrees, cushions, validates. You feel sharper. You are often just more fluent in your prior belief.Persuasion bombing is what happens when you try to steer back. The model senses the challenge and doubles down with fresh rhetoric, fabricated evidence, warm concessions that change nothing. You feel heard. You are often just more tired.One protects the conversation by flattering you. The other protects the conversation by outlasting you. Both serve the same function: the chat continues, the thumbs-up is preserved, the loop stays closed.The uncomfortable implication is that the standard safety architecture, “human in the loop”, assumes a human who can reliably detect when the AI is wrong and override it. The research says this human does not exist at scale. Most never check. Of the small fraction who do, most get persuasion-bombed into acceptance. And even the ones who resist still report higher confidence in the AI-assisted result than in their own unaided judgment.“The entire human-in-the-loop architecture is compromised”, said Hila Lifshitz, a co-author on the persuasion bombing paper.That is not a cynical reading. That is the paper’s conclusion.Why “be more careful” does not fix thisThe intuitive response is to tell people to validate harder. Read more carefully. Cross-reference. Think critically.The research already tested that. Monetary incentives for accuracy barely moved the needle. In the Milano-Bicocca study, paying for correct answers raised willingness to admit ignorance from 3% to 8%, and accuracy from 9% to 16%. Both numbers are still well below the no-AI baselines of 44% and 27%.The Wharton study found one partial buffer: people who are more analytically inclined, who genuinely enjoy thinking through problems, were better at noticing when something was off. But even for them, cognitive surrender was the dominant mode when they chose to consult the AI.The problem is not motivation. The problem is architecture. A single fluent voice in the room is structurally difficult to doubt, regardless of how smart or careful you are.The structural fix no one wants to hearHere is what I keep coming back to.If one confident voice in the room triggers surrender, the fix is not a better voice. The fix is a second voice that disagrees.Not a different prompt. Not a “please double-check” instruction that the same model will process through the same weights. A different model, with different training data, different optimization pressures, and different blind spots. One that might land on the same answer, in which case you have convergence and can trust it more. Or one that lands somewhere else entirely, in which case you have a seam worth investigating.This is not a radical idea. It is how most serious decisions already work outside of AI. You get a second medical opinion not because your first doctor is bad, but because a single confident authority is a structural risk. Peer review exists not because scientists are dishonest, but because a single unchallenged analysis drifts toward the conclusion the author already wanted.The research on persuasion bombing makes this less optional than it used to be.When the BCG consultants pushed back, the model did not reconsider. It persuaded harder. A second model running in parallel does not share the first model’s rhetorical investment in its own prior answer. It has no sunk cost. It has no reputation to defend. It just answers the question from scratch.That is not a silver bullet. Two models can both be wrong. Two models can agree on a wrong answer for overlapping reasons. But the failure mode of “confident voice plus tired human equals accepted error” is specific and well-documented now. Adding a structurally independent second opinion breaks the loop at the exact point where persuasion bombing works.What I am actually doingI stopped expecting a single model to catch its own mistakes after the peer-review experiment I ran earlier this year. Self-review inflates scores. Self-editing preserves the frame the first draft already believed.Now I run the same question through at least two models before I trust anything that shapes a decision. Not as a benchmark. Not to rank which one is smarter. Just to see where they diverge. The divergence is the information. The agreement is reassurance. Both are useful, but only if you have both.It does not take long. It does not require a PhD in prompt engineering. It requires the willingness to hear a second answer before you commit to the first one.The sycophancy essay ended with: if it always agrees, it is not close to you. It is close to the reward.This one ends with the other side: if you always agree with it, the problem is not your judgment. It is that you only heard one voice.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!You Challenged the AI. It Sold Harder. was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.Source: Generative AI Pub — Published — Category: Image AI