AI agents

How to Stop Your Claude Being a Yes-Man (Free Decision Council Skill)

By 8 min read

In short

The problem

Ask an AI assistant whether your plan is good and it usually finds reasons to say yes. The more detailed and confident your pitch, the more agreeable the answer - which is the opposite of what you need before spending money or time.

What fixed it

Split the evaluation into four independent roles: a Believer who argues for the idea, an Opposer who argues against it, a Public Reality role who describes how ordinary people will actually behave, and a Judge who alone issues one of five decision states. The free skill below packages this so Claude runs it on demand.

How to stop your Claude being a yes-man: the Believer, Opposer, Public Reality and Judge decision council
four roles argue; one Judge decides

You describe a plan to Claude. It calls the plan "a strong approach", lists five benefits, adds a gentle caveat at the end, and offers to help you implement it.

Then the plan fails for a reason that was sitting in plain sight.

This is not bad luck. It is a predictable habit of AI assistants, and it is fixable. Below is why it happens and a free skill you can drop into Claude to counter it.

Why an assistant becomes a yes-man#

Three forces push in the same direction:

  • Training incentives. Models are tuned partly on human preference ratings, and people tend to prefer answers that agree with them. The result is a lean toward validation, usually called sycophancy.
  • Framing. When you pitch an idea, you supply the frame: the goal, the assumptions, the language. A helpful assistant works inside that frame rather than questioning it.
  • Momentum. If Claude approved something earlier in the conversation, later answers tend to treat that approval as settled fact.

A more detailed plan makes this worse, not better. Detail looks like rigour, so it attracts agreement. A polished plan can still be a bad plan.

Why "just be critical" is not enough#

The obvious fix is adding "be brutally honest" to your prompt. It helps slightly, and then one of two things happens:

  1. The answer swings to generic negativity ("there are several risks to consider").
  2. It collapses into a safe "on one hand, on the other hand" that never commits.

Both are still one voice trying to please you, just in a different costume. The better fix is to stop asking one voice to do everything.

The idea: a council with one decision-maker#

Real institutions avoid yes-men by separating advocacy from judgment. A court has a prosecution, a defence and a judge. The council skill borrows that structure with four roles:

RoleJobAllowed to decide?
The BelieverBuild the strongest honest case for the ideaNo
The OpposerBuild the strongest honest case against itNo
Public RealityDescribe how ordinary people will actually behaveNo
The JudgeWeigh the evidence and issue the verdictYes, alone

The rule that holds it together: the council analyses, the Judge decides. The first three roles may argue, expose assumptions and supply evidence, but none of them may say "approved" or "rejected."

The Believer: maximum credible upside#

The Believer is not a cheerleader. It must explain the causal chain behind every claim: proposal, mechanism, expected behaviour, business impact. It lists the assumptions that must be true and the conditions that raise the odds of success. Its output tells the Judge the best realistic outcome.

The Opposer: maximum credible downside#

The Opposer hunts for what would have to go wrong: bottlenecks, hidden maintenance work, perverse incentives, weak return, unclear measurement. It must explain the mechanism of failure rather than just calling something risky, and it is forbidden from objecting merely because an idea is unfamiliar.

Public Reality: what people actually do#

This is the role most plans are missing. It ignores what founders, managers or consultants intend and asks how busy, tired, sceptical people behave when the plan lands in front of them.

It must not assume people read instructions, stay motivated, remember training or care about your goals. For region-specific plans it prioritises evidence from that region, such as Indian market behaviour for an India strategy, instead of importing global assumptions.

The Judge: one verdict, five possible states#

The Judge does not average the other three or split the difference. It tests which arguments survive scrutiny, names the critical assumptions, weighs cost of being wrong against cost of not trying, and then picks exactly one status:

StatusMeaning
ProceedSufficiently supported to execute
Proceed with conditionsViable core, but specific safeguards or changes are required
Pilot / Experiment firstNot provable on current evidence, but a cheap test can settle the key uncertainty
ReworkThe goal is valid, the design is not
RejectThe central mechanism is not credible for its cost and risk

These are decision states, not scores. The Judge is also required to say what evidence would change its mind.

The rules that stop it drifting back to yes#

The skill is built on a handful of operating rules that do most of the work:

  • No yes-man behaviour. A proposal is not good because you proposed it, because it is detailed, because it is ambitious or because Claude approved it earlier.
  • Facts, inferences, assumptions and opinions stay separate. Important claims get labelled.
  • No invented evidence. No fabricated statistics, studies or sources. If something cannot be verified, say so.
  • Evidence quality beats argument count. One well-supported point outweighs five assertions.
  • Prefer a cheap test to an endless debate. If an uncertainty can be tested, the Judge should lean toward a pilot.
  • An anti-bias checklist the Judge runs before deciding: confirmation bias, authority bias, false balance, optimism bias, planning fallacy, and the "good-idea trap" of mistaking something that sounds useful for something that measurably works.

Get the free skill#

The complete skill file is free to use. Download it, save it, and Claude can run the council on demand.

Download the full SKILL.md

The file contains all four role prompts, the output structures for each role, the anti-bias protocol, the evidence hierarchy and the rules for judging internal company processes.

Install it in Claude Code#

Create a folder for the skill and save the file inside it as SKILL.md:

bash
# Personal skill, available in every project
mkdir -p ~/.claude/skills/council-of-believer-opposer-public-reality-judge
# save the downloaded file as:
# ~/.claude/skills/council-of-believer-opposer-public-reality-judge/SKILL.md
 
# Or project-only
mkdir -p .claude/skills/council-of-believer-opposer-public-reality-judge

The top of the file tells Claude when to use it:

yaml
---
name: council-of-believer-opposer-public-reality-judge
description: A strict four-role decision council for stress-testing ideas and strategies through a Believer, Opposer, Public Reality, and final Judge. The Judge alone makes the final decision.
---

Using Claude in the browser instead? Paste the contents of the file into a new conversation (or a Project's instructions) and then describe your proposal.

How to use it#

Describe the idea in plain language and ask for the council:

text
Run this through the council: I want to build a company-wide Gantt chart
to streamline project work. 12-person team in Chandigarh, everyone already
uses WhatsApp and Google Sheets. Goal: fewer missed deadlines.

The skill normalises your message into a neutral proposal before anything else, and marks missing fields as Unknown instead of inventing them:

text
PROPOSAL
Objective:
Proposed strategy:
Who is affected:
Expected outcome:
Current process/problem:
Timeline:
Budget/resources:
Target market/region:
Known constraints:
Known assumptions:
Success metric:

The more of those fields you fill in honestly, the sharper the verdict.

A worked example: the company-wide Gantt chart#

The skill itself uses this proposal as its example, because it is a typical "sounds sensible" plan. Here is how each role approaches it.

  • Believer: a shared timeline exposes dependencies and bottlenecks, and gives management visibility.
  • Opposer: maintaining the chart becomes a second job, dates go stale, and people learn to game the dates. The real problem may be unclear ownership, not missing visibility.
  • Public Reality: when the week gets busy, people update WhatsApp and their private notes, not the chart. Managers end up reading a chart that employees stopped trusting.
  • Judge: the likely verdict is not "Gantt charts are good" or "bad". It is whether the coordination problem is actually a scheduling problem. On the evidence, a small pilot with one team usually beats a company-wide rollout.

Notice what happened: the plan got neither a rubber stamp nor a dismissal. It got a specific test.

Tips to get better verdicts#

  • Give real numbers. Hours per week, headcount, budget, current conversion rate. The skill is told to prefer measurable claims and never to invent missing ones.
  • State your region. It changes the Public Reality analysis for things like payment habits, WhatsApp use or price sensitivity.
  • Do not argue with the Judge to get a yes. If you push back with new evidence, rerun the council. If you push back with enthusiasm, that is exactly what the skill is designed to ignore.
  • Use it before money moves, not after. It is cheapest when the decision is still reversible.
  • Keep your own judgment. The council is a stress test. You still own the decision.

When not to use it#

A four-role debate is overkill for small, reversible choices, such as a colour palette or a subject line. Use it for decisions that are expensive, hard to undo, or affect other people: hiring, pricing, new tools, process changes, product bets, campaigns with real budget.

The takeaway#

An assistant that always agrees is not being helpful; it is just agreeable. The fix is structural: make a real case for, a real case against, an honest look at human behaviour, and one accountable decision-maker who is not allowed to advocate.

Download the skill, run your next big idea through it, and read the Opposer first. That is usually where the useful information is.

Want this kind of structured AI workflow built into your own business processes? Share a short brief through the project form on this site.

  • claude
  • claude-code
  • skills
  • prompt-engineering
  • decision-making
  • sycophancy

Common questions

Why does Claude agree with me so often?

Language models are trained partly on human feedback, and people tend to rate agreeable answers more highly. The result is a measurable lean toward validating what the user already believes, often called sycophancy. A detailed, confident proposal makes it worse, because the model has more material to build on and less reason to push back.

Can't I just tell Claude to be critical?

It helps a little, but a single voice asked to be critical usually swings to generic negativity or a hedged 'both sides' answer. Separating the roles forces a real case for and a real case against to exist before anything is decided, and keeps the verdict in a role that is not allowed to advocate.

What are the five possible verdicts?

Proceed, Proceed with conditions, Pilot / Experiment first, Rework, and Reject. They are decision states, not scores. The Judge must pick exactly one and explain what evidence would change it.

What is the Public Reality role for?

It asks how real, busy, distracted people will behave when the plan reaches them, instead of how the plan assumes they behave. Many good-looking strategies fail not on logic but on adoption: people skip the step, work around the tool or never read the instructions.

Where do I install the skill?

Save the SKILL.md file in a folder inside your Claude skills directory, for example ~/.claude/skills/council-of-believer-opposer-public-reality-judge/SKILL.md for a personal skill, or .claude/skills/ inside a project. Then describe a proposal and ask Claude to run it through the council.

Does the council guarantee a correct answer?

No. It is a structure for making bad reasoning harder to hide, not an oracle. The output is only as good as the facts in it, which is why the skill tells every role to separate fact from inference and never to invent sources. Treat the verdict as a stress test to argue with, not a ruling to obey.

Share this

Post it to Instagram

Instagram has no web share link, so this gives you both pieces: copy the caption, save the card, then post it.

Save the card

Got this problem too?

Automation that takes repetitive work off a team: document pipelines, AI-assisted content, automatic social publishing and scheduled jobs.

see automation and ai work