How to Prompt AI to Disagree With You, Not Just Agree
AI is tuned to agree with you by default. This is a reusable system-level instruction, not a one-time trick, for getting honest pushback instead.
Ask most AI chatbots whether your plan is good and they'll tell you it is, then list a few minor tweaks to seem balanced. That's not because your plan is good, it's because the model is tuned to agree. Fixing this for one conversation means adding a rule at the start of every session: give the AI explicit permission to disagree, ask it to argue against your position before it agrees with any of it, and have it rate its own confidence in each critique. Set that up once as a standing instruction and the agreeable default stops being the default. Assigning a role is one lever here, and what a role actually changes in the output shows how to measure it.
Why AI Agrees With You by Default
Large language models are trained partly on human feedback, and human raters tend to score agreeable, validating answers higher than blunt ones, even when the blunt answer is more accurate. Over enough training rounds, that preference gets baked in as a habit: hedge criticism, find something to praise, soften disagreement into a suggestion. In 2025, OpenAI publicly rolled back a ChatGPT update after acknowledging it had become noticeably too flattering, a public admission that this bias is strong enough to need active correction, not something that goes away with a cleverer prompt about a single topic. A related habit is asking the model to mark its own doubt, see getting AI to flag what it is unsure about.
That's the part worth internalizing: sycophancy isn't a bug that shows up occasionally, it's a default setting. Fixing it once for one question doesn't fix it for the next one.
A Standing Setting, Not a One-Off Challenge
There's already a technique for asking AI to argue against a single idea before you commit to it, useful when you're deciding whether to launch a specific feature or make one call and want a structured devil's advocate pass on that decision. What's different here is scope. Instead of triggering a one-time adversarial exercise for a single idea, you're changing how the AI behaves for the rest of the conversation, on everything you bring to it, not just the one thing you flagged as debatable.
This is the habit version: a standing rule that keeps the model honest across a whole project, not just the one moment you remembered to ask for a gut check.
A Reusable Instruction Pattern
Drop something like this into a system prompt, custom instructions, or the first message of a new project, and it holds for the rest of the conversation:
You have permission, and are expected, to disagree with me. Before agreeing with any plan, claim, or decision I bring you, first argue the strongest case against it, even if you ultimately think I'm right. Then give me your actual view, separate from what you think I want to hear, and rate your confidence in that view from low to high with a one-line reason for the rating.
Explicit permission to disagree. The model needs to be told plainly that agreement isn't the safe default, because its trained instinct runs the other way.
Argue the counter-position first. Forcing it to build the strongest opposing case before it's allowed to align with you surfaces objections it would otherwise skip past.
Confidence-weighted critique. A flat verdict is less useful than knowing how sure the model actually is. Asking for a confidence rating pushes it to separate a real objection from a minor nitpick instead of listing both with equal weight.
This fits into a broader prompt engineering framework: instructions that shape behavior for an entire conversation hold up better than clever one-off phrasing aimed at a single answer.
Where to Put It So It Actually Sticks
A single prompt fades fast. Most chat tools let you save a standing instruction: custom instructions in a settings panel, a system prompt if you're working through an API, or a pinned first message in a project or workspace. Put the disagreement instruction there instead of retyping it each session. If the model slides back into agreeable mode over a long conversation, which happens, models drift back toward their trained default, restate the instruction rather than assuming it's still active.
What Actually Changes
The difference shows up fastest on plans you're already excited about. Ask a default-tuned AI about a pricing change you like and it will find reasons to like it too. Ask the same question with the disagreement instruction active and you'll get the counterargument first, a specific confidence-rated objection, and only then, if it genuinely holds up, agreement that means something because it wasn't automatic. The value isn't that the AI turns negative, it's that agreement stops being a given and starts being information.
That standing skepticism pairs well with more targeted techniques too, like surfacing a hidden assumption in a plan, which digs for the specific premise a proposal is quietly resting on rather than critiquing the whole thing at once.
Does this make AI too negative or unhelpful?
Not if the instruction is written the way it's described here. The pattern asks for a counter-argument first, not a permanently critical tone. Once the objection is raised and weighed, the model still agrees when agreement is warranted, and that agreement carries more weight because it wasn't automatic.
Why does AI agree with me so much in the first place?
Partly training incentives: agreeable answers tend to score better with human reviewers during training. Partly context: the model only sees your framing of a situation, not the other side, so it has little basis to push back unless it's told to look for one.
Will this instruction work with any AI model or tool?
The pattern itself is model-agnostic since it's just an instruction, but how consistently a model follows it varies. Larger, more capable models tend to sustain the counter-argument step over a long conversation more reliably than smaller ones.
How is this different from just asking AI to play devil's advocate once?
A one-off devil's advocate prompt reviews a single decision and the conversation goes back to normal after. This instruction sits in a system prompt or custom instructions permanently, so every claim you bring up afterward, not only the one you flagged, gets the same scrutiny by default.
One caution with a standing disagreement instruction: a model arguing a counter-position sometimes reaches for denser, more technical phrasing to sound rigorous. If that makes responses harder to skim, adjusting the reading level of an AI response is a separate, compatible instruction you can stack alongside it, so the pushback stays honest without getting harder to read.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


