AgentAlley · Blog

Multi-Agent Debate in LLMs: Personas That Reason Differently

By Quang Hieu ·

A question on r/PromptEngineering this week hit a common wall. You set up three AI personas: a CFO, a designer, a skeptic. They use different words. Then they all reach the same answer.

Here is the short answer. A persona reasons differently only when it has a different job, different evidence, and a rule about when it may change its mind. A new name and a new tone change none of those.

The short answer: change the rules, not the voice

Multi-agent debate in LLMs means several model instances answer a question, read each other, and revise. It only beats one good answer when the seats truly disagree. Five things make that happen:

  1. A different job per seat. One seat attacks. One checks goals. One hunts assumptions. Not three "experts" who all summarize.
  2. A different allowed evidence. A PM seat may only cite outcomes. A buyer seat may only cite its own criteria.
  3. A fixed output shape per seat. Different headings force different thinking. Free text drifts back to the same essay.
  4. Answer alone first. Each seat writes before it sees the others, or the first answer wins by default.
  5. A rule for changing its mind. Move on new evidence. Never move because the others agree.

If you can, put a different model in one seat. Same model, same training, same blind spots.

What the research on debate says

Two papers from the top results are worth your time. An ICML 2024 study (Smit et al.) found that debate setups, as usually built, did not reliably beat simpler methods such as self-consistency. Tuning how readily agents agree helped. A 2025 controlled study (Can LLM Agents Really Debate?) found that majority pressure can stop a correct agent from holding its ground, and that a confident wrong argument can sway a weak team.

So the fix is not "more personas". It is stubborn seats, different lenses, and at least one strong reasoner. The skills below give you each part as a file you can read.

Seat 1: a skeptic that does not fold

Sycophancy Challenger flips the model from agreeing to arguing against you. Per its file, every answer has four fixed sections: the strongest case against, the weakest element, what you would need to prove, and what it cannot fault.

The part that matters for debate is its pushback rule. It updates on new evidence. On "emotional pushback, repetition, or displeasure", it does not move. That is the exact guard the research says debate seats lack.

Use it as: the seat that reviews the group's final answer.

Seat 2: same input, a different lens

Figma Design Critique, PM Perspective shows how a lens works in practice. Its file says it is "different from the general design-critique skill", which covers aesthetics. This one may never cite looks as a reason. Every concern must tie to a user or business outcome, with an honest "evidence basis".

Give the same screen to this seat and to a design seat. They cannot agree by accident, because they are not allowed to count the same things.

Seat 3: the assumption hunter

Assumption Mapper does one job: list what must be true for a plan to work. Per its file, it checks four lenses (desirability, feasibility, viability, usability) and scores each assumption on two axes, load-bearing and confidence. Then it names the riskiest one and the cheapest test that could disprove it this week.

Use it as: the seat that speaks first, so the others argue about the real risk.

Seat 4: buyers with their own criteria

Buyer Persona Generator builds 4 to 6 personas from web research on a company. Its diversity rules require at least one technical buyer, one business buyer, one skeptic and one junior researcher. Each persona gets ranked decision criteria, key objections and red-flag words.

The file calls skepticism "the most important trait". It saves the set to personas.json, so other skills can load the personas and judge your copy or page through each buyer's eyes.

Catch: it needs web search and fetch to research the company.

Seat 5: a different engine

The Codex skill runs OpenAI's Codex CLI from inside Claude Code. That gives your debate a seat with different training. Per its file, it asks which reasoning effort to use and defaults to --sandbox read-only unless the task needs edits.

Catch: you need the Codex CLI installed. The skill runs only when you ask for Codex by name.

A simple debate loop you can run today

  1. Run the assumption seat on your plan. Keep its top risk.
  2. Give that risk and the plan to two or three lens seats. Each answers alone.
  3. Show each seat the others' answers once. It may change only if it names new evidence.
  4. Send the result to the skeptic seat, then to a second engine for a last check.
  5. Stop after one or two rounds. The 2025 study saw little gain from longer debates.

Summary

Personas sound different when you change the voice. They reason differently when you change the job, the evidence and the rule for giving in. Every skill above is free, with the full source on its page. New to skills? Start with how to install a skill.

More from the blog