Two tools, one prompt, and a sushi buffet
I gave two AI tools the same design prompt.
Claude behaved like a UI designer sitting beside me, especially when I was working in Figma, a popular design app. Codex behaved like a designer, a developer, a researcher and a slightly unhinged creative director sharing one trench coat. Claude also went through my credits like it had arrived at an all-you-can-eat sushi buffet.
So which one was better? Wrong question. If I wanted a design partner inside Figma, one answer. If I wanted range and initiative, another. If I cared about my credits, that was a third conversation entirely.
I kept landing in the same place across code, video, architecture and creative work: there is no permanently best AI. There is a bench.
During a discussion about software architecture, I noticed the silly public AI tests say the same thing. Having an AI run a vending machine, play an old video game, or draw a pelican riding a bicycle reveals things tidy leaderboard scores miss.
Why “which AI is best?” has no useful answer
A leaderboard ranks AI models either by standard tests (math puzzles, exam questions, coding tasks) or by which answer people preferred in quick side-by-side votes. Either way, it tells you who did best on someone else's questions, not which one writes the best reply to your customers or stays inside your budget.
Think of hiring. The candidate with the best exam scores might still be wrong for your front desk; you'd rather watch them handle a real afternoon. AI is the same. A model that wins a reasoning test can still stumble on your forms, fail to recover when a step breaks, or forget a rule after a few weeks.
That's why I like the silly tests. Each one exposes something a short quiz can't:
- A vending machine run over a long task shows whether the AI can plan and adapt as things change.
- An old video game shows whether it can see what's on screen, remember what happened, and recover from a mistake.
- A pelican on a bicycle, or a plate of spaghetti, shows whether it understands how physical things fit together.
Test scores reward neat questions with one right answer. Votes reward the answer that looked best at a glance on someone else's prompt. Your work is messier. (For the fun version of AI being confidently wrong, see The Infinite Confidence Tournament.)
Build a bench from your own work
Your bench is a handful of real, typical tasks where you already know what a good answer looks like. That last part matters most. If you can't tell a good answer from a bad one, you can't score anything.
Say you help run a small bakery. Your bench might be three tasks:
- Turn an annoyed customer email into a kind, firm reply.
- Turn this week's specials into three short social posts.
- Summarize a supplier's price list and point out what went up.
For each task, write a short checklist before you try any tool. For the customer reply: apologizes once, doesn't promise a refund we don't offer, names the fix, under 120 words, sounds like us. Now “good” isn't a feeling. It's five things you can tick. If you've tried the practice loop in AI does the work faster. Judging it is still your job, this is the same yes/no checklist, now pointed at two tools instead of your own attempts.
Three to five tasks is plenty. Include one tricky one; that's where tools differ most.
Compare fairly: same input, same checklist, names hidden
Every tool gets the same input, the same background information, and roughly the same time or budget. Change the prompt between tools and you're testing your prompts, not the tools.
Score against your checklist. I also watch for:
- Accuracy. Did it make anything up?
- Editability. How much do I have to fix before I'd use it?
- Initiative. Did it notice something I didn't ask about?
- Cost and speed. What did it burn to get there?
- Failure. When it went wrong, did it admit it or bluff?
Don't score only the first answer. Count the back-and-forth it took to reach something usable.
If you can, score blind: paste the answers into a document as “A” and “B”, or have a friend strip the names. Brand names nudge your opinion more than you'd think.
A scored customer reply (invented)
- Apologizes once: A yes, B yes
- No refund we don't offer: A yes, B no (“we'll refund you in full”)
- Names the fix: A yes, B yes
- Under 120 words: A no, B yes
- Sounds like us: A yes, B no
A scores 4 of 5, B scores 3. B also promised money the bakery doesn't offer, which rules it out for this task. No leaderboard can tell you that.
Give tools roles, not crowns
The result of a good bench is rarely one winner. It's a set of roles: this one drafts, that one checks. Useful roles include researcher, drafter, critic, visual generator, long-document reader, and fast cheap rewriter.
Roles let tools work as a team: one drafts, another critiques, you decide. The value comes from different perspectives, not an expensive group chat where five models politely agree.
Write your findings as one plain sentence per task: For this kind of task, with this information and budget, this tool currently does best. That's less exciting than “Model X destroys everything.” It's also useful.
Notice “currently.” Models and prices change quickly, so re-run the same bench when something new arrives. For a longer walkthrough, try the guide Compare AI tools on your own tasks.
Try it
Try it in 20 minutes
You need two chat assistants open in two browser tabs (for example ChatGPT, Claude, Gemini or Copilot) and a blank document. Use invented material only, never real customer or private information.
- Start with two ready-made tasks
Copy the two tasks below, then add a third task like your own work, with invented details. It will tell you the most.
You work for a small bakery. We don't give refunds, but we can offer a free replacement cake or store credit. Reply to this customer: “My birthday cake was meant to arrive Saturday and came Sunday. The party is over. What are you going to do about it?”
Suggest five nut-free weeknight dinners that each take under 30 minutes.
- Write what good looks like first
Copy these checks, and write four or five of your own for task 3 before you run anything. Reply: apologizes once; doesn't promise a refund; offers one specific fix; under 120 words; sounds like a person. Dinners: exactly five; every dish is actually nut-free (check sauces, pesto and toppings yourself); each takes under 30 minutes; no repeats.
- Give both assistants the exact same prompt
Paste identical text into each tab. Don't tweak the wording for one of them.
- Shuffle, then score
For each task, flip a coin to decide which tab's answer goes under A and which under B, and write the key at the very bottom of the document. Paste every task's answers first, then score them all in one pass without looking at the key. It won't be perfectly blind, but it keeps the brand names out of sight while you judge. Note how many follow-ups each needed.
- Assign roles, not a winner
Write one sentence per task: “For [task], I'll start with [tool] because [reason].” Then save the tasks and checklists so you can re-run them when tools change.
You’ll know it worked when
You have one page with two or three tasks, checklists, scores, and a sentence saying which tool you'll use for what. If the better tool changed from task to task, the bench did its job.
When you’re ready
Going further: testing AI for a team
At work, the bench grows into three layers:
- Small test cases: short examples with known answers and a clear checklist, like your personal bench.
- Realistic run-throughs: multi-step tasks with missing information or a step that fails, to see whether the AI recovers or bluffs.
- Running beside the people doing the job: the AI works on the same real cases but takes no action, and you compare its results with theirs before trusting it.
Include ugly cases: missing details, contradictory instructions, out-of-date documents, text inside a file that tries to give the AI orders, and a request it should refuse or hand to a human.
Track success, unsupported claims, time people spend correcting it, cost, and whether it answers the same way twice. Every real failure becomes a new test case, with private details removed. The goal isn't a perfect score. It's knowing where the tool is safe to use, and what happens when it steps outside that zone.
Avoid these
Common mistakes
- Choosing by leaderboard or viral demo
Those show how a model did on someone else's test. Five minutes with your own task tells you more.
- Testing tasks you can't judge
If you don't know what a good answer looks like, every answer looks fine. Build your bench from work you know well.
- Only judging the first answer
The first reply is a demo. Count how much back-and-forth it takes to reach something you'd actually use.
- Crowning a permanent winner
Tools and prices change often. Keep your bench saved and re-run it when something new arrives.
Copy, adapt, run
Prompts to try
Paste one into any chat assistant and replace anything in [brackets].
Bench builder
Help me build a small test set for comparing AI tools on my own work. I regularly do these tasks: [list 3 to 5 tasks]. For each one, write one realistic example input using invented details (no real names or private data) and a checklist of 4 to 6 yes/no items describing what a good answer must include or avoid. Keep the checklist short enough to score in a minute.
Team evaluation planner
I'm considering using AI for this workflow at work: [workflow]. Identify who uses the result, what decisions depend on it, what a mistake would cost, what information it can use, and where a human must approve. Then propose five short test cases with known answers, two realistic run-throughs that include missing information or a failed step, and a way to run it beside the people doing the job without it taking any action. Connect every measure to a real risk or outcome.
The tool kit
Tools and links
- Any two chat assistants (ChatGPT, Claude, Gemini, Copilot)All the 20-minute bench needs.
- Compare AI tools on your own tasksThe step-by-step guide version of this bench.
- Anthropic test and evaluation guidePlain guidance on defining success and building test cases.
- OpenAI evaluation guideFor teams: how to test prompts and models systematically.
- OpenAI EvalsOpen-source test framework, for builders.
- MLflow GenAI evaluationEvaluation and monitoring for teams running AI in production.
The short version
What to remember
- “Which AI is best?” has no useful answer. “Which AI is best at this task, for me, right now?” does.
- A bench is a few of your real tasks where you already know what a good answer looks like.
- Test fairly: same input, same checklist, names hidden, and count the follow-ups.
- Give tools roles instead of crowns, and re-run the bench when tools or prices change.
patrickz
