Run the build, send a screenshot, describe the chaos
At 11 p.m. I was stress-testing a game I was rebuilding from my childhood. I’d built it in Godot, a game engine I had never opened before this project. The game was a mess. The camera got stuck behind walls. Characters spawned inside walls. Some fights never ended, so the whole game froze in place (gamers call that a softlock). Textures looked wrong.
My job settled into a simple loop: run the build, send a screenshot, describe the chaos.
Then something changed. The model stopped acting like autocomplete and started acting like a debugging partner. It traced likely causes, then began building its own small debugging tools: labels, markers and logs that showed things I couldn’t see. Once the game made cause and effect visible, the AI had something real to reason about.
This wasn’t the AI building the game on its own. It was a different split of the work. I owned the goal, the taste and the final call on whether something was fixed. The AI dug through code, came up with explanations and built the tools to test them.
I didn’t become a Godot expert that night. I learned something I can use anywhere: AI is only as good as the evidence in front of it.
Why “it’s broken” gets you a confident wrong answer
Hand a chat assistant a vague complaint and it fills the gaps with the most likely story, told in a calm, sure voice. Sometimes the story is right. Often it’s plausible and wrong, and it sounds the same either way. (Why it sounds so sure: How AI actually answers.) Speed is not certainty.
Picture a household spreadsheet where the monthly total looks too low. You type “my total is wrong, fix it.” The AI might blame rounding, a hidden row or a number stored as text. It picks one and rewrites your formula. Now you have a formula you don’t understand and the same wrong total. Nobody looked at anything. You swapped guesses.
The fix isn’t a smarter AI. It’s better evidence. Debugging means finding out why something doesn’t do what you expected, and the method is the same for a video game, a spreadsheet, a prompt that keeps giving bad answers or a smart-home routine that turns the lights on at 3 a.m.
The evidence loop in six steps
- Reproduce it. Find the shortest steps that cause the problem every time. “Sometimes the total is off” is a rumor. “When I add a row under Water, the total doesn’t change” is a bug you can work on.
- Say what you expected and what actually happened. “Expected 58.00, got 51.50” gives the AI a specific gap to explain. “It’s off” doesn’t.
- Make the invisible visible. Most problems hide in something you can’t see. More on this below.
- Ask for a few ranked hypotheses, and how to kill each one. A hypothesis is a possible cause. Ask for two or three, ranked by the evidence, each with one check that would prove it wrong. A guess you can disprove is useful. A guess you can only believe is not.
- Make the smallest change. One fix, aimed at the confirmed cause. A big rewrite while you’re unsure just buries the cause.
- Retest the original problem and something nearby. Run your exact steps again, then try a neighbor: a new row, a slightly different question. Fixes like to break the thing next door.
Before you change anything, keep a safe checkpoint: a copy of the spreadsheet, the old prompt saved in a note or, for code, a snapshot in Git (a tool that keeps versions of your code).
Make the invisible visible
This is the step that turned my Godot night around. A camera stuck behind a wall doesn’t tell you why. The camera problem needed a view of where the camera was aiming and of the collision layers (the invisible shapes that decide what bumps into what). Spawning inside walls needed a marker on every spawn point and a check for clear space. Endless fights needed a log of each change in the fight’s state and a way to reset one fight without restarting the game.
None of those tools were the fix. They were temporary windows, easy to delete afterward. And the AI can build them for you. That’s the part most people miss: AI isn’t only good at guessing answers. It’s good at making small tools that show whether a guess is true.
You don’t need code for this. In a spreadsheet, add a helper column showing the running total row by row, so you can see where the sum stops counting. For a prompt that keeps giving bad answers, ask the AI to list the assumptions it made before answering. For a smart-home routine that misfires, check its activity history or have it notify you each time it runs. Same move every time: turn “it’s acting weird” into something you can point at.
Try it
Try it in 15 minutes
You need any chat assistant. A spreadsheet app is optional. If you have one, type the six items into B2:B7 and put the step 3 formula in B8 so you can check things yourself. Don’t show the AI the formula until step 3.
- Report only what you can see
Paste the prompt below. It gives what you see and what you expected, but not the formula.
My club snack budget has six items in column B, rows 2 to 7: Chips 12.50, Juice 18.00, Fruit 9.75, Cookies 7.25, Napkins 4.00, Water 6.50. The total in B8 shows 51.50, but my calculator says 58.00. Give me up to three possible causes, ranked, each with one check that would prove it wrong. Do not fix anything yet.
- Notice the confident guesses
The AI may spot that exactly 6.50 (Water) is missing, but it can’t see why. The range could stop early, or Water could be stored as text. Both fit the numbers. Notice how sure it sounds anyway.
- Make the invisible visible
The hidden detail is the formula. In a spreadsheet, click B8 and read the formula bar, or turn on Show Formulas (Formulas tab in Excel; View › Show › Formulas in Google Sheets). Then tell the AI the formula.
The formula in B8 is =SUM(B2:B6).
- Watch guesses become evidence
Ask the question below. Several possible stories should narrow to one cause you can point at.
Which of your causes does this confirm, and which does it rule out?
- Fix small, then test next door yourself
Accept the smallest change to B8. Then test it instead of asking. Insert a row just above the total, add Plates 5.00, and see whether the total becomes 63.00. No spreadsheet? Work it out yourself: the fixed range ends at B7, and Plates now sits in row 8. Is row 8 inside the range? If Plates isn’t counted, you found the nearby weakness the first fix missed. Paste that result back and ask for a fix that survives new rows.
You’ll know it worked when
You can name the cause, the one piece of evidence that ruled out the other guesses, and whether your fix survives a new row, because you checked, not because the AI said so.
When you’re ready
If you write code: bring richer evidence
With an AI coding agent (one that can read and edit your project’s files) the loop is the same, with richer evidence: a screenshot, exact steps, relevant log lines, the scene or file name, recent changes and what you’ve tried. Treat context like a budget. Too little and the agent guesses. Too much and it drowns.
Ask for the plan before the edit. An agent that jumps straight to changing code may fix the symptom and hide the cause. Practice with the Debug with evidence guide, then try the character debug lab. If you’re building something new rather than fixing it, build in small, proven slices so there’s less to debug.
Avoid these
Common mistakes
- Letting it fix before you can reproduce
If you can’t make the problem happen on purpose, you can’t tell whether a fix worked. Get repeatable steps first.
- Believing the most confident answer
A sure tone is the AI’s default, not a sign it’s right. Ask what evidence would prove the cause wrong, then go look.
- Changing five things at once
If it works afterward, you don’t know which change mattered. If it breaks, you don’t know which change broke it. One change, then retest.
- Only retesting the exact bug
The original steps pass, and the fix quietly broke the thing next to it. Always try one nearby case, and keep a checkpoint.
Copy, adapt, run
Prompts to try
Paste one into any chat assistant and replace anything in [brackets].
Bug report (no fixes yet)
Help me figure out why [spreadsheet / prompt / routine / program] isn’t working. Expected: [result]. Actual: [result]. Steps that cause it every time: [steps]. Recent changes: [list]. Evidence: [error text, screenshot description or log lines]. Give me at most three possible causes, ranked by the evidence. For each, name one observation that would confirm it and one that would prove it wrong. Do not fix anything yet.
Make it visible
Suggest the smallest temporary way to see what’s happening inside [the formula / the routine / the scene]: a helper column, a log, a label or a marker. It must not change how the thing works, must be easy to remove, and must help me tell your top two hypotheses apart.
Fix and retest
Make the smallest fix for the cause we confirmed. Tell me how to re-run the original steps plus these nearby checks: [list]. List what you changed, what’s still uncertain and how to remove any temporary debug tools.
The tool kit
Tools and links
- GitSafe checkpoints you can roll back to (for code).
- Godot EngineThe free game engine from my story.
- Godot debugging toolsGodot’s built-in ways to see inside a running game.
- CodexOr another coding agent that can see your project.
The short version
What to remember
- Make the problem happen on purpose before you change anything.
- Give the AI evidence: expected, actual, exact steps and what you can see.
- Ask for ranked hypotheses and what would prove each one wrong.
- Make the smallest fix, retest nearby and keep a checkpoint. Speed is not certainty.
patrickz
