The Council

All guides

Compare GPT, Claude, Qwen and Gemini on the same question

Side-by-side tabs do not compare AI models fairly. Put GPT, Claude, Qwen and Gemini under the same rules, with sealed first answers, and see where they differ.

The trouble with side-by-side tabs

The usual way to compare models is to paste one prompt into several chat tabs and read the answers. It tells you something, but less than it seems. Each answer is a first draft that nobody challenged. You cannot tell a firm view from a guess, and you never learn what one model would say to another’s best point.

The same rules for every model

On The Council every seat gets the same standing instructions, the same word limit and the same clock. The page writes the instructions, so you do not have to keep four prompts in step.

A fair comparison in five steps

  1. Ask one question that has a decision in it. “Which of these two designs should we build?” compares models better than “tell me about these designs”.
  2. Seat each model once, with no role. Roles are for testing an idea. To compare the models themselves, let each speak for itself.
  3. Keep the openings sealed. Unsealed, the second speaker has read the first, and you are comparing one answer with one reply to it.
  4. Use a fixed length. A word limit stops a model from looking thorough by writing more.
  5. Write the motion yourself. Then every model votes on your wording, not on a summary one of them drafted.

For a quick comparison set the rounds of debate to zero and switch off the closing statements. You get the sealed openings, the final vote and the chair’s statement. Add rounds when you want to see how the answers hold up under challenge.

Reading the comparison

After the vote the page shows each member’s score as a dot on one line from 0 to 10. Dots that bunch together mean the models ended close. Dots at opposite ends mean they did not, and the list underneath tells you what each one holds.

In the minutes, look at “What moved”. A model that changed its position for a stated reason has told you more than one that only repeated itself. So has a model that held its ground and said exactly why.

Two cautions. The page prints the first with every result.

Comparing two models from the same provider

Two seats can use the same account with different models. Seat a provider’s small model and its large one, put the same question, and you will see whether the larger one changes the outcome or only the price.

Which models, and what it costs to run them

A seat connects with an API key from OpenAI, Anthropic, Alibaba Model Studio (Qwen), Google AI Studio (Gemini), DeepSeek or OpenRouter, or to a model running on your own machine. Usage is billed by the provider, to your own account.

If you only have chat subscriptions, seat the models by hand. The page writes each message, you paste it into ChatGPT, Claude or Qwen Chat, and you paste the answer back. It is slower, and it works with any model that has a chat box. There is a guide to debating by hand, without an API key.

Your keys and your sessions stay in your browser.

Seat your models or read how a moderated AI debate works.

Open the council

Updated