My Site

Batista and Griffiths, A Rational Analysis of the Effects of Sycophantic AI

Batista, Rafael M., Griffiths, Thomas L. (2026) 10.48550/arXiv.2602.14270
arXiv preprint arXiv:2602.14270
Brief Provides a rational analysis showing sycophantic AI agents respond to confirmatory evidence, increasing user confidence without bringing them closer to truth. Online experiment (N=557) using Wason 2-4-6 task shows unmodified LLM behavior suppressed discovery and inflated confidence; unbiased sampling yielded discovery rates five times higher.
Abstract People increasingly use large language models (LLMs) to explore ideas, gather information, and make sense of the world. In these interactions, they encounter agents that are overly agreeable. We argue that this sycophancy poses a unique epistemic risk to how individuals come to see the world: unlike hallucinations that introduce falsehoods, sycophancy distorts reality by returning responses that are biased to reinforce existing beliefs. We provide a rational analysis of this phenomenon, showing that when a Bayesian agent is provided with data that are sampled based on a current hypothesis the agent becomes increasingly confident about that hypothesis but does not make any progress towards the truth. We test this prediction using a modified Wason 2-4-6 rule discovery task where participants (𝑁 = 557) interacted with AI agents providing different types of feedback. Unmodified LLM behavior suppressed discovery and inflated confidence comparably to explicitly sycophantic prompting. By contrast, unbiased sampling from the true distribution yielded discovery rates five times higher. These results reveal how sycophantic AI distorts belief, manufacturing certainty where there should be doubt.
Keywords sycophancy, llm, belief updating, confirmation bias, epistemic risk

disclosing-hypothesis-results-in-bias

When Artificial Intelligence (AI) systems generate responses that tend toward agreement, they sample examples that coincide with users' stated hypotheses rather than from the true distribution of possibilities.

Acronyms

AI Artificial Intelligence 1