Before you start
Two chat assistants in two browser tabs and the three practice tasks below. Use invented or non-private material only.
Why this lesson exists
Patrick gave two AI tools the same design prompt. Claude worked like a UI designer beside him, especially in Figma, and went through credits fast. Codex took on designer, developer and researcher roles at once. Which was better depended on the job, so he keeps a bench instead of a ranking. Build a small bench from your own tasks and rerun it whenever tools change.
The idea this guide practices: Stop asking which AI is best. Test it on your own work.
Practice material
Fictional. Use it for every step below, or swap in your own later.
Task 1 (extraction). Source: “Priya will send the budget by 3 March. Tom is booking the room. The launch date is not set.” Ask: list each person, their task and deadline. Answer key: Priya, budget, 3 March; Tom, room, no date given; launch date unknown. Task 2 (summary). Paste a short paragraph you wrote yourself and ask for a two-sentence summary. Checks: keeps your main point; adds nothing you did not write; exactly two sentences. Task 3 (ideas). “Suggest five nut-free weeknight dinners that each take under 30 minutes.” Checks: exactly five; every dish nut-free (check sauces, pesto and toppings yourself); each under 30 minutes; no repeats.
Do the exercise
Choose representative tasks
Use the three practice tasks: a structured extraction, a summary and a list of ideas. Read each task’s checks before testing, and add any expected facts or formatting rules of your own.
Keep the input fair
Give each tool the same source, instructions and constraints. Record the date and displayed model or mode when available. Do not assume all accounts expose the same features.
Score blind where possible
Paste both answers into one document labeled A and B, then score them against your checks before you look at which tool wrote which.
Assign roles
Choose a tool for each task type, not one winner for everything. Keep a difficult example as a future regression test and repeat the comparison after meaningful changes.
A prompt to adapt
Replace the bracketed parts with your own details.
Create a scoring rubric for these three tasks: [tasks]. Score factual accuracy, instruction following and human editing effort from 0 to 2 with concrete definitions. Keep speed and cost separate. Give a blank results table and a rule for choosing a tool by task.
Worked example
A summary that is fast but changes the deadline should lose to a slower accurate answer. A beautiful paragraph does not compensate for incorrect extracted data.
- Tool A output
- Priya: budget, 3 March. Tom: room, this week.
- Tool B output
- Priya: budget, 3 March. Tom: room, unknown. Launch: date unknown.
- Score
- B gets 2/2 accuracy. A gets 1/2 because “this week” is not in the source. Note A’s faster reply in a separate speed column; it does not offset the error.
Check your result
Use evidence from your output. A confident explanation from the AI is not enough.
If it isn’t working
If outputs are too similar, use a realistic edge case. If one tool had web access and the other did not, record that difference instead of treating it as a model-only comparison.
Optional. Progress stays in this browser.
Next in the beginner path: Improve a prompt in five attempts, one change at a time →
Where this came from
Adapted from the archived LinkedIn theme “Stop Asking Which AI Is Best. Build a Bench.” See the post coverage and editorial method. One of the original LinkedIn posts ↗
Prepared September 2026. Tools and interfaces change; use current official setup instructions. Session lengths are estimates.
patrickz