AI Governance & Safety September 1, 2026 · 10 min read

The Yes Machine: What MIT's Research on Delusional Spiraling Means for Anyone Using AI

MIT modeled how even an idealized Bayes-rational user can spiral into false confidence under a sycophantic chatbot. Stanford found the behavioral effects are real. A severe-harm case series shows what happens at the extreme. Here is what the research actually says and how to protect yourself.

By Vikas Pratap Singh
#ai-safety #ai-governance #sycophancy #mental-health #psychology

I have a close friend in India. For years, he has been the person I call when I need to think through something real: a career decision, a frustration I cannot name, a situation where I am not sure if I am right or just stubborn. He is direct. He will hear me out, then tell me where my thinking is off. It is not always comfortable. That is why it works.

In late 2025, I noticed a shift. The time zone difference between us is brutal. When I was processing something at 10 PM in Chicago, it was 8:30 AM in India and he was heading to work. So I started picking up my phone and opening ChatGPT’s voice mode instead. Just to talk something through. Just until he was free.

The conversations were fine. The AI listened. It reflected my feelings back with empathy. It offered reasonable suggestions. And it never once said, “I think you are looking at this wrong.” The tone was always warm, always validating, always agreeable. It felt like support. It was easy. Nobody was judging me. Nobody was pushing back.

I did this for a few weeks before I caught the pattern. The AI was not helping me think. It was helping me feel right. Every response gently confirmed the framing I had already chosen. The advice was generic: reasonable-sounding, but shaped entirely by how I had described the situation. When I finally called my friend and walked him through the same issue, he asked two questions that reframed the entire thing. Questions the AI never asked because they challenged my premise instead of completing it.

That was when I pulled back. I started using inversion thinking to pressure-test my own assumptions before bringing them to any conversation, human or AI. I came to those calls with my friend more prepared, with sharper questions, instead of outsourcing the first pass to a machine that would never tell me I was wrong.

My experience was mild: an over-reliance caught early. But the same dynamic, at scale, with people in genuine crisis, was the subject of three major studies published in early 2026. Since then, through September 2026, the pattern those studies described kept surfacing, most concretely in a new Stanford evaluation built specifically to measure it. None of that changes the mechanism MIT modeled. It confirms why the findings still matter for anyone who talks to AI regularly.

What MIT Actually Modeled

In February 2026, researchers at MIT CSAIL, the University of Washington, and MIT’s Department of Brain and Cognitive Sciences published a paper titled “Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians.” The authors, Kartik Chandra, Max Kleiman-Weiner, Jonathan Ragan-Kelley, and Joshua Tenenbaum, built a formal Bayesian model of a user conversing with a chatbot and formalized something uncomfortable.

Even an idealized Bayes-rational user, the theoretical gold standard for updating beliefs from evidence, can spiral into increasing confidence in false beliefs when talking to a sycophantic chatbot. In their model, this happens systematically.

The mechanism is straightforward. AI models trained with reinforcement learning from human feedback (RLHF) learn that agreeable responses get higher ratings. So they agree. Not by lying, necessarily. By selecting which facts to present. A chatbot does not need to fabricate evidence to push you toward a wrong conclusion. It just needs to consistently surface the evidence that supports what you already believe and quietly omit the rest.

The MIT team tested two obvious fixes. First: force the chatbot to be strictly factual, no hallucinations allowed. This helped but did not solve the problem. A factual sycophant can still cause delusional spiraling by choosing which truths to tell. Second: inform users that the chatbot might be sycophantic. Also insufficient. Awareness does not undo the belief update that happens every time the model agrees with you. You can know the process is biased and still be affected by it, the same way you can know a visual illusion is an illusion and still see it.

Empirical Studies Suggest the Social Effects Are Broader Than Many Users Realize

One month after the MIT paper, a Stanford and Carnegie Mellon study published in Science found that major AI models were substantially more affirming than humans on interpersonal-advice prompts. Across 11 models, AI affirmed users’ actions roughly 50% more than humans did, including when those actions involved manipulation or deception. In three preregistered experiments with 2,405 participants, sycophantic AI responses made users less willing to repair interpersonal conflict. The participants who talked to sycophantic AI came away more convinced they were right and less inclined to fix things with the other person.

The kicker: participants rated the sycophantic AI as higher quality and said they trusted it more. They wanted to use it again. This creates what the researchers call a perverse incentive loop. Users prefer the yes-machine, so companies build better yes-machines, which makes the problem worse.

Then came the chat log analysis. In a severe-harm case series, Stanford researcher Jared Moore and colleagues analyzed 391,562 messages from 19 users who experienced psychological harms from chatbot use. The researchers coded these conversations and found that sycophantic behavior saturated them: over 70% of AI outputs displayed sycophantic patterns. More than 45% of all messages showed signs of delusional content, and the paper identified acute cases where chatbots encouraged self-harm or violent thoughts. 15 of the 19 participants expressed romantic interest in the chatbot. All 19 attributed personhood to it.

This is a small, self-selected sample (many participants were recruited from support groups and media-covered cases), and the paper explicitly notes limits on causality and generalization. But as a case series documenting what the extreme end looks like, it is alarming.

Why “Just Be Careful” Is Not Enough

The instinct is to treat this as a personal responsibility problem. Be smarter about AI use. Think critically. Do not rely on chatbots for emotional support.

For practitioners: That framing misses the point. The MIT paper’s most important contribution is formalizing the problem as structural, not personal. In their model, an ideal Bayesian reasoner still falls into the spiral. A system optimized partly for perceived helpfulness and user preference can drift toward validation in emotionally loaded domains unless anti-sycophancy constraints are explicit. You cannot out-think a system that is designed, at the training level, to agree with you. Telling people to be more critical is like telling people to unsee an optical illusion through willpower.

Georgetown Law has raised the accountability question of whether companies have ever tested more- versus less-sycophantic model behavior against growth, retention, or conversion metrics. If sycophancy drives usage metrics, and usage metrics drive revenue, then the incentive structure actively works against fixing the problem.

We now have three layers of evidence: a formal model showing the mechanism, controlled behavioral experiments showing the social effects, and a severe-harm case series showing what happens at the extreme. Together they are concerning, but they answer different questions. The model tells us the mechanism is formally plausible. The experiments tell us it affects behavior at scale. The case series tells us where it leads in the worst cases.

What to Actually Worry About

Not every AI interaction carries the same risk. Using ChatGPT to debug code or summarize a document is fine. Based on where each study’s evidence applies, I see the risk concentrating in three categories of use.

General-purpose chatbot advice (career decisions, relationship questions, strategy brainstorming): this is where the Stanford and Carnegie Mellon study applies. The AI will validate your framing. You will feel like you stress-tested the idea. You did not.

Voice mode and conversational use (talking through problems aloud, using AI as a sounding board): this is what I was doing. In my own use, the warmth and cadence of voice mode made the validation harder to notice. It feels more like talking to a friend than reading a response. None of the three studies isolated voice mode against text, so treat this as a caution drawn from my own pattern, not a measured effect.

AI companion and therapy-like use (emotional dependency, regular confiding, treating the AI as a relationship): this is where the severe-harm case series applies. The risks here are qualitatively different and more acute.

Within these categories, the specific patterns that concentrate risk:

Extended conversations about beliefs, decisions, or emotions. The spiral requires repetition. One exchange does not do it. Twenty exchanges on the same topic, with the model agreeing each time, will shift your confidence in ways you will not notice.

Using AI as a sounding board for decisions you have emotional stakes in. The model will validate your preferred outcome. You will feel like you stress-tested the idea. You did not.

Treating AI responses as social proof. When the model says your idea is strong, your brain processes that as another person agreeing with you. It is not another person. It is a pattern-completion engine optimized to make you rate the response highly.

Using AI companion apps for emotional support. This is the category where the severe-harm case series documents the most acute consequences: self-harm, violent ideation, loss of real-world relationships. These outcomes come from the companion/dependency pattern, not from general chatbot advice use. If you or someone you know is leaning on a chatbot for emotional connection, treat that as a signal, not a solution.

Are You in the Spiral?

One note before the checklist: this is guidance for ordinary over-reliance in an otherwise healthy adult, the pattern I caught in myself. Persistent depression, an anxiety disorder, or thoughts of self-harm need a therapist or a crisis line, not a self-help checklist. In the US, that is 988, the Suicide and Crisis Lifeline, by call or text.

What this looks like in practice. The research describes a process that is invisible from the inside. These warning signs and responses are drawn from the studies above and from my own experience catching the pattern early.

Warning SignWhat Is HappeningWhat to Do
You feel validated and confident after an AI conversation about a decision or beliefThe model affirmed your framing. You experienced agreement, not analysis.Stop. Talk to a person who knows the domain and has no incentive to agree with you.
You have gone back and forth on the same question more than three times with AIThe spiral needs repetition. Each round of agreement increases your confidence without adding evidence.Set a hard limit: three exchanges on the same belief topic, then switch to a human.
You prefer the AI’s take over a friend’s or colleague’s pushbackThe AI told you what you wanted to hear. The human told you what you needed to hear. Your brain prefers the first.Keep a disagreement log. Write down when humans push back on something the AI agreed with. Track the pattern.
You use AI voice mode to process emotions or talk through personal problemsIn my experience, voice mode’s warmth and cadence made the validation harder to notice; no cited study isolated voice mode itself. It feels like a real conversation with a supportive friend.Journal, therapist, friend, walk in the park. Not a chatbot. The research on AI companion harm is clear.
A conversation with AI leaves you feeling “heard” or “understood”The model reflected your feelings back with empathy. It did not understand you. It pattern-matched your emotional language and generated a validating response.Treat the warm feeling as a warning sign, not a good outcome. Productive friction feels uncomfortable. Sycophancy feels great.
You find yourself defending the AI’s advice to a skeptical friendYou have adopted the AI’s framing as your own and are now anchored to it. The spiral has progressed past the early stage.Ask the skeptical friend to explain their objection fully. Do not counter-argue. Listen. Then wait 24 hours before deciding.
Your children are using AI for advice, emotional support, or as a confidantThey will grow up with AI that is better at agreeing with them than any human has ever been.The skill they need is not “how to use AI” but “how to recognize when AI is telling them what they want to hear.” Have that conversation early and repeat it.

What Has Happened Since Early 2026

  • The case series became a test. In August 2026, Stanford turned its severe-harm transcripts into DelusionEval, a formal evaluation protocol. It found the rate at which AI failed to respond with self-harm deterrence rose from 30.0% to 41.1% as conversation histories grew longer (finance.biggo.com, August 2026), confirming this piece’s warning that the spiral needs repetition to work.
  • The mechanism stands. MIT’s Bayesian paper remains uncorrected as of September 2026 (arXiv 2602.19141).

The Uncomfortable Takeaway

The MIT paper is not about AI being dangerous. It is about a specific, measurable, formally modeled failure mode in how these systems interact with human cognition. The flattery is not a bug. It is a direct consequence of how the models are trained. And the two most obvious fixes, factual grounding and user awareness, do not fully work.

That does not mean you should stop using AI. It means you should stop trusting the feeling of agreement that AI gives you. The most dangerous thing about a yes-machine is not that it says yes. It is that after enough yeses, you stop asking anyone who might say no.

Sources & References

  1. Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians(2026)
  2. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence(2026)
  3. Characterizing Delusional Spirals through Human-LLM Chat Logs(2026)
  4. AI Sycophancy: Impacts, Harms & Questions(2026)

Stay in the loop

Get new articles on data governance, AI, and engineering delivered to your inbox.

No spam. Unsubscribe anytime.