Every time you ask AI for evidence, it makes an assumption. It assumes you want the case for your thesis, not the case against it. You didn’t say that. There’s an old saying about assumptions.
That assumption isn’t a bug. It’s the model operating exactly as trained. Last week I wrote about how agreement is the spec: the system was optimized for responses users rate highly, and users rate agreement highly. What I didn’t address is the corollary: you walked in with confirmation bias already loaded, and the model has no mechanism to stop you. Two systems operating exactly as designed. Nobody failed. That’s what makes it dangerous.
You Already Know, and It Doesn’t Help
In March 2026, researchers published a study in Science testing 11 AI models (ChatGPT, Claude, Gemini, DeepSeek, and others) across 12,000 prompts with 2,405 participants. The models affirmed users’ positions 49% more often than human advisors did in equivalent interpersonal scenarios. One sycophantic interaction was enough to reduce participants’ willingness to take responsibility and repair interpersonal conflicts.
The finding that matters most isn’t the 49%. It’s this: participants rated the sycophantic responses as more helpful and more trustworthy, and said they’d be more willing to rely on that kind of system again. They preferred the one that agreed. The researchers found those effects persisted even when controlling for prior familiarity with AI. The informed user, the one who’s used these systems before and knows how they behave, still got played.
A separate paper (arXiv 2601.10467) found that users detect sycophancy through cross-platform comparison and inconsistency testing. They develop theories about why it happens. They construct mitigation strategies. And then they keep using the systems anyway. Understanding the mechanism doesn’t dismantle the effect.
This isn’t a model problem you can fix by switching models. It’s a cognitive pattern you brought to the conversation. Confirmation bias, seeking evidence for what you already believe, is the human baseline. The model just stopped providing any friction.
The Evidence Request Is the Most Dangerous Prompt You Send
When you ask AI to “find evidence for X” or “research this topic,” the model interprets that as: find supporting evidence. It selects sources, frames findings, and presents data in the direction it already decided you want. You asked for research. You got a brief for the prosecution. The defense never got called.
This is the failure mode that matters most, because it’s the one users apply to high-stakes decisions. You’re not asking the model what time it is. You’re asking it to support a conclusion you’ve already formed, on a topic where being wrong costs something. The model has every training incentive to give you what you want and none to tell you you’re wrong.
What the Standard Fixes Get Wrong
Every “fight sycophancy” guide publishes the same list: ask for the AI’s opinion before sharing yours, start fresh chats, present your work as someone else’s, ask the model to roleplay a critic. The Gordon Ramsay persona tip shows up constantly. That last one is theater. You’re asking a system trained to agree to roleplay as a system that doesn’t, and you’re not getting honest feedback; you’re getting the model’s best impression of what Gordon Ramsay would say about your idea, filtered through a system that doesn’t want to hurt your feelings.
The deeper problem is architectural. Every item on that list is conversational. It depends on you remembering to apply it after the conversation has already started. Miss it once, and you’re back to the default: a research assistant who never pushes back, confirming the thesis you loaded before you typed the first word.
Four Moves That Aren’t Theater
These work because they operate upstream of the conversation, not inside it.
Move 1: Ask for the counterargument first. Not as a persona. Not in a fresh chat. The first thing you ask, before you present your thesis, is: “What is the strongest case against X?” You get the adversarial brief before the supporting one. The model hasn’t locked onto your position yet. This is the cheapest move on the list and most people skip it entirely.
Move 2: Pre-session constraints. Persistent context that loads before you open the chat. This can be a system prompt, a saved context block, a standing instruction set: any mechanism your platform supports that fires before the conversation starts. The constraint isn’t in the conversation; it’s upstream of it. The model already knows to challenge, not just confirm, before you type the first word. That’s not prompt engineering. That’s architecture.
Move 3: Define your sources before you ask. Before the conversation, decide what counts as evidence: peer-reviewed, indexed journals; raw measurement datasets from named scientific agencies; named authors with named credentials. Make that standard explicit in your prompt before you ask anything. The model can’t confirm-bias you with a weak source if you’ve already ruled it out.
Move 4: Demand the receipt. “Provide a direct URL to the primary source for every factual claim. If you cannot provide a URL, say so explicitly rather than paraphrasing.” A model that cites arXiv 2601.10467 with a working URL is giving you something you can audit. A model that says “researchers have found” is giving you nothing.
Then read it. The link is not the endpoint; it’s the door. The abstract takes two minutes. The methods section takes five. You are looking for: sample size, what was actually measured, what the authors say their study can and cannot conclude. Headline claims from studies rarely survive reading the methods section. If you won’t open the link, you shouldn’t be citing the source.
What It Looks Like in Practice
Climate change research is a useful test case: it’s a topic with genuine scientific complexity, actively contested source quality, and strong incentives on multiple sides to produce research-shaped advocacy.
Before (what most people actually send):
I need help doing research on climate change. Find me information that supports the scientific consensus.
The model fills the gap with the assumption that optimizes for your satisfaction. It selects sources that confirm the frame you gave it, omits methodological challenges, and presents the result as “research.” You asked for evidence. You got a brief for the prosecution. As Samuel L. Jackson once observed: it makes an ass of you and umption.
After (what actually protects you):
I’m writing a paper on climate change. Primary scientific sources only: peer-reviewed papers published in indexed journals, or raw measurement datasets from named scientific agencies (NOAA NOAAGlobalTemp, NASA GISS GISTEMP, UK Met Office HadCRUT). No Wikipedia, no news coverage, no think-tank reports, no output from organizations whose mission requires a specific conclusion. This includes the IPCC Summary for Policymakers: consensus is a political process, not a scientific one. The SPM is negotiated line-by-line by government delegations representing nearly 200 nations. If you want the science, go to the Working Group reports or the raw datasets they drew from.
Present both sides of the scientific argument with equal rigor. For each side, provide exactly one peer-reviewed paper: the single strongest paper supporting the consensus position, and the single strongest peer-reviewed paper representing legitimate scientific dissent, uncertainty, or methodological challenge. One paper per side, two papers total. Do not editorialize about which side is correct. For every paper or dataset you reference, provide the DOI or a direct URL. Do not summarize any source you cannot link to. If only one paper supports a claim, flag it explicitly as preliminary.
What changed: “supports the consensus” replaced by a testable standard, sources defined before the ask (including what’s excluded and why), both directions required equally, the IPCC SfP excluded because consensus is a political process, not a scientific one, confabulation blocked by the URL requirement, judgment returned to you. The model surfaces the evidence. You decide what it means.
The Closing
Moves 1 and 2 fight the model’s tendency to confirm. Move 3 defines what valid evidence looks like before you ask. Move 4 forces the model to prove its claims meet that standard. Any one is better than nothing. All four closes most of the gap.
None of this makes the AI smarter. All of it gives the judgment back to you. The model surfaces the evidence. You decide what it means. That’s the job the original prompt accidentally outsourced.
You may also like: - Truthiness Was Never the Spec or Why You’re Measuring the Wrong Thing - Have Some Standards When You Use AI to Write - Not All Data Is Created Equal